DocsQuick StartAI News
AI News<think>**Translating leaderboard achievement** </think> Wan3.0 Breaks into the Top Three on the Arena Video Leaderboard for the First Time
New Model

<think>**Translating leaderboard achievement** </think> Wan3.0 Breaks into the Top Three on the Arena Video Leaderboard for the First Time

2026-09-07T12:07:09.428Z
<think>**Translating leaderboard achievement**

</think>

Wan3.0 Breaks into the Top Three on the Arena Video Leaderboard for the First Time

In the Arena blind test from August 31 to September 6, Alibaba’s Wan3.0 debuted at No. 3 on the image-to-video leaderboard, achieving an ELO score of 1,481—53 points higher than its predecessor, Wan2.7. Its advantage lies not only in more visually stunning output, but also in its growing ability to compete for production-grade video workflows through improved stability and continuity.

<think>Planning markdown-preserving translation

</think>

Wan3.0 Enters the Top Three of Arena’s Video Rankings for the First Time, as Alibaba Bets on “Stable Production”

Alibaba’s Wan3.0 entered the Arena video model rankings for the first time—and immediately claimed a top-three spot.

The rankings for Week 36 of 2026, published by the Arena platform, show that in anonymous blind tests conducted from August 31 to September 6, Wan3.0 ranked third in the Image-to-Video category, with an ELO score of 1,481 and a win rate of approximately 57%. Ahead of it were MiniMax-H3 and Gemini Omni 1.1 Flash. Wan3.0 was 7 points behind the second-place model and 16 points behind the leader.

The significance of this result does not lie in the conclusion that “a Chinese model has finally entered the top three.” Rather, it lies in the clear gap between Wan3.0 and its predecessor: Wan2.7 previously ranked tenth, with an ELO score of 1,428, while Wan3.0 improved by 53 points in a single release.

For a video model, this is not an ordinary version iteration. It means Alibaba is moving the Wan series from “able to generate a good-looking video” toward “able to continuously generate usable assets across multiple shots and under multiple conditions.” The latter goal is more difficult—and much closer to the real needs of advertising, short-form drama, music videos, and film previsualization.

Illustration showing Wan3.0 ranking in the top three of Arena’s video model rankings, labeled with its score of 1,481, win rate of 57%, and score difference from its predecessor Wan2.7

A Top-Three Blind-Test Result Shows That User Preferences Are Changing

Arena’s rankings are based on anonymous blind-test votes. Users typically cannot see the model names. Instead, they compare outputs from different models generated from the same prompt or reference material, and then choose the one they prefer. The platform uses an ELO system to calculate rankings, so the results are closer to “which model users are willing to choose in an actual comparison” than to standardized test scores covering every capability dimension.

This is particularly important for video models.

Video generation does not have a single clearly defined answer like a mathematics problem. Image quality, naturalness of motion, subject consistency, cinematography, style matching, and prompt adherence can all influence a user’s judgment at the same time. A video with an exceptionally impressive opening frame may still receive a poor evaluation if the character’s face changes in the second shot, the clothing shifts, or the motion trajectory suddenly loses physical logic.

Therefore, Wan3.0’s entry into the top three indicates at least one thing: in the subjective choices of real users, it is no longer winning merely through single-frame clarity or a particular stylized effect. Instead, its overall viewing experience is approaching that of the current first tier.

Of course, a score of 1,481 should not be directly interpreted as meaning that “Wan3.0 has comprehensively surpassed all overseas models.” Different Arena sub-rankings are independent of one another, and ranking scores are also affected by sample size, prompt types, and voter preferences. When score differences are small, ranking changes may fall within the range of statistical error. New models also need to accumulate a sufficient number of votes before moving from the preliminary stage into a more stable ranking.

In other words, this is a strong signal, but not a final acceptance report. It proves that Wan3.0 deserves serious comparison, but it cannot replace stress testing in a company’s specific workflow.

The Core Upgrade in Wan3.0: From Short Clips to Continuous Production

Wan3.0 officially launched on August 24. Alibaba positions it not as a single-purpose text-to-video model, but as an “all-purpose reference video generation model.” According to the model description on Alibaba Cloud Bailian, wan3.0-video supports multimodal inputs including text, images, video, audio, and file links, covering the following scenarios:

  • Text-to-video;
  • Image-to-video, including first-frame and first-and-last-frame control;
  • Reference-to-video generation;
  • Generation of characters, styles, props, and scenes driven by reference materials;
  • 480P, 720P, and 1080P output;
  • Video generation of up to 30 seconds per request.

Thirty seconds alone is not a parameter sufficient to make a model a leader. The real challenge is that, within those 30 seconds, the model must handle multiple stages of action, changes in shots, and relationships between subjects.

For example, a car-chase sequence might include a character getting into a vehicle, the vehicle starting, a road chase, a high-speed turn, and an explosive finale. If the model can only generate attractive images at each individual moment, it still cannot complete a usable shot. It needs to make the characters, vehicles, environment, and direction of motion connect coherently along the timeline.

This is also where video generation models have most commonly broken down in the past: the first second looks excellent, the body begins to deform by the third second, the prop changes appearance by the fifth, and by the eighth second, the character and the scene no longer belong to the same setting. Creators have had to repeatedly generate alternatives, extract usable segments, and then hide flaws through editing.

Wan3.0 is attempting to solve precisely this problem: “Every frame looks right individually, but the sequence looks wrong when connected.”

Multimodal Input Changes More Than Just the Interaction Method

Another noteworthy design choice in Wan3.0 is its support, for the first time, for document formats such as doc, xls, ppt, pdf, and md.

At first glance, this may look like nothing more than adding several file types to the input box. In practice, however, it changes the upstream video-generation workflow. In the past, video models typically received a prompt, an image, or a reference video. To provide a model with product materials, course outlines, or storyboard scripts, companies often had to manually reorganize them into prompts before generating individual shots.

If a model can directly read structured documents, the input expands from “a single description” to “an entire set of creative materials.” A product presentation can provide brand colors, product selling points, page structure, and marketing tone. A PDF can provide setting details and character backgrounds. An Excel spreadsheet might contain product information, a shot list, or a subtitle table.

This does not mean that uploading a PowerPoint will automatically produce an advertisement requiring no revisions. Document parsing, shot planning, image generation, and post-production editing are still separate stages, and the model may not always accurately understand complex layouts or implicit intentions. But for standardized content production, reducing the need for manual translation between formats is itself an efficiency gain.

More importantly, document input makes it easier for video models to connect to existing enterprise systems. Brand teams do not have to learn an entirely different creative language from scratch. Materials originally used for sales, training, and product explanations can all potentially become part of a video-production asset library.

“Realism” Is No Longer Just About Resolution and Detail

This time, Alibaba placed particular emphasis on the realism of people, the reproduction of physical properties, and stylistic consistency. Three words repeatedly mentioned in testing feedback were: stable, realistic, and textured.

Here, realism does not simply mean increasing the resolution. The problems that most easily expose AI-generated people are often not insufficiently clear pores, but the lack of a coherent relationship between facial details, emotions, and movement: the eyes look in one direction while the head faces another; the corners of the mouth are smiling, but the shoulders and arms show no corresponding emotional change; the skin appears delicate, yet the overall result has a plastic quality.

Wan3.0 attempts to establish more stable coordination among facial features, skin, micro-expressions, and body movements. For short-form drama, advertising, and character-driven narratives, this is more important than simply pursuing an image that “looks more like a photograph,” because audiences judge whether a character feels natural based on continuous movement and emotional changes, not on a single screenshot.

Its handling of stylized reference images is also worth noting. Traditional Chinese illustration, steampunk, stage drama, and high-saturation advertising visuals all require the model to maintain consistency in materials, colors, and composition. If every shot change alters the lighting, clothing textures, and character proportions, a stylized video will quickly lose its overall coherence.

From a practical perspective, Wan3.0’s strength is more about “reducing rework” than “generating a masterpiece every time.” This is also what distinguishes it from models used on social media to showcase a single spectacular shot: what enterprises actually pay for is usually not an occasional success, but a stable supply of editable, revisable, and deliverable assets.

API Pricing Is Not High, but Costs Cannot Be Judged Only by the Per-Second Rate

wan3.0-video currently provides inference services through Alibaba Cloud Bailian. Its publicly listed prices are:

| Output Resolution | Price | | --- | ---: | | 480P | RMB 0.3/second | | 720P | RMB 0.6/second | | 1080P | RMB 1.2/second |

For a single 30-second video, the original prices work out to approximately RMB 9, RMB 18, and RMB 36, respectively. During the public beta, Alibaba Cloud Bailian and the Qwen AI platform offered limited-time discounts. Specific promotional prices and validity periods should be confirmed on the console.

These prices are relatively competitive in the video model market, especially for batch generation of storyboards, advertising assets, and social media short videos. However, developers should not compare model costs solely on the basis of “price per second.”

What really needs to be calculated is the total cost of a deliverable video:

  1. What is the first-pass generation success rate?
  2. How consistent is the same character across multiple generations?
  3. How many retries are required to correct hands, text, or props?
  4. Does a 720P generation require additional upscaling?
  5. Can a long video be used directly, or must it be divided into multiple clips and reassembled?
  6. Will queue times, concurrency limits, and manual post-generation review slow down the overall workflow?

If a model has a low per-unit price but requires more than a dozen generations to produce a usable version, its actual cost may be higher than that of a seemingly more expensive model with a higher first-pass success rate. The significance of Wan3.0 entering Arena’s top three is also that it may reduce these hidden costs.

What Developers Should Actually Test

If you are considering integrating Wan3.0 into a product, it is not advisable to use only a few cinematic prompts for a demonstration. A more effective evaluation should focus on failure points in the actual business workflow.

1. Test Subject Consistency

Prepare front, side, and full-body reference images of the same character, then generate different scenes continuously. Observe whether the hairstyle, clothing, facial features, and body shape remain stable. This is a fundamental metric for short-form drama and virtual-avatar products.

2. Test Long-Sequence Actions

Do not test only walking and turning around. Include actions such as getting into a vehicle, picking up an object, opening a door, running, fighting, and interacting with multiple people. Observe whether the actions contain skipped frames, object interpenetration, or breaks in causality.

3. Test Adherence to Reference Materials

Provide clear color palettes, composition, and style references, and require the model to maintain consistency across multiple shots. In particular, test both realistic and stylized materials, as they place different demands on the model’s control capabilities.

4. Test Structured File Input

Upload real product presentations, course handouts, or storyboard tables instead of simple files prepared specifically for demonstrations. Focus on whether the model can extract key information, understand page hierarchy, and convert the materials into a coherent shot plan.

5. Test Text and Brand Elements

If the video contains packaging, logos, subtitles, or product names, verify text accuracy separately. Video models remain prone to errors with dynamic text, complex fonts, and small brand marks. Do not overlook this simply because the overall image looks realistic.

Alibaba’s Entry into the Top Three Means the Competition Is Moving from “Can Generate” to “Can Deliver”

At present, overseas models still dominate Arena’s overall rankings, while Chinese models are more concentrated in specialized capabilities. A similar shift has also appeared in the front-end development rankings: Alibaba’s Qwen3.8-Max-0902 entered the rankings for the first time at fourth place, while Qwen3.8-Flash-Next, Tencent’s HY4-Preview, and Zhipu’s GLM-5.3-Flash simultaneously entered the top fifteen.

This shows that the competitive path for Chinese models is becoming more specific. Rather than catching up across all general-purpose metrics, it may be more effective to deepen capabilities in vertical scenarios such as video, coding, Chinese-language knowledge, and enterprise workflows. Wan3.0’s ranking is one validation of this strategy in the video field.

But the top three is only the beginning. The next challenge for video models is not how to generate another more attractive demo, but how to maintain stability with more complex inputs, higher concurrency, and stricter copyright and content-safety requirements. Creators need controllable characters, reusable shots, predictable costs, and interfaces that can connect to existing editing and asset-management systems.

If Wan3.0 can translate the user preferences reflected in Arena into success rates and delivery efficiency in production environments, it will have a genuine chance of becoming part of the video infrastructure. Otherwise, its top-three finish will remain merely an impressive blind-test result.

For developers, the most sensible approach is not to immediately migrate the entire existing video pipeline, but to first conduct a comparative test using their own materials: the same character, the same script, and the same resolution. Compare Wan3.0 with existing models in terms of success rate, number of retries, and post-production repair time.

Rankings tell you which models are worth trying; production data determines which ones are worth keeping.

Sources

The information in this article is current as of September 7, 2026. Arena’s rankings are based on anonymous blind-test voting and ELO calculations, reflecting subjective user preferences rather than absolute capability tests. Specific model prices, rate limits, and service availability should be verified against the latest information in the official console.

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: