DocsQuick StartAI News
AI NewsWan 3.0 Public Beta: Turn Documents into 30-Second Videos
New Model

Wan 3.0 Public Beta: Turn Documents into 30-Second Videos

2026-08-06T15:04:37.461Z
Wan 3.0 Public Beta: Turn Documents into 30-Second Videos

Alibaba is launching the Wan 3.0 public beta tonight, increasing the maximum video length per generation to 30 seconds and adding support for document inputs such as PDFs, PowerPoint presentations, and spreadsheets for the first time. What truly deserves attention is not the longer duration, but that video generation is beginning to move into workplace content production.

Alibaba Cloud announced this evening (August 6) that its next-generation video generation model, Wan3.0, has officially entered public beta. The two most notable changes in this upgrade are the ability to generate videos up to 30 seconds long in a single run and, for the first time, support for document inputs such as PDFs, PowerPoint presentations, Word documents, spreadsheets, and Markdown files, in addition to text, images, audio, and video.

Wan3.0 is now available in public beta on Alibaba Cloud Model Studio, Wanjing Yike, the official Wanxiang website, and the PC version of Qwen Creation. The Qwen app is being made available through a phased rollout. Alibaba has also announced API pricing, although the API has not yet been opened to all users. According to the company, full availability is coming soon.

Wan3.0 public beta interface and illustrations of 30-second video generation and document input features

The Significance of 30 Seconds Goes Beyond Simply Making Clips Longer

Wan3.0 can generate a single 30-second video clip in one run. Compared with the few-second outputs commonly seen today, this may sound like a simple increase in duration, but it actually touches on one of the hardest problems for video models: consistency over a long timeline.

A five-second shot may be considered successful if it completes a single action, such as a person turning their head, a car driving past, or the camera moving forward. At 30 seconds, however, the model must handle many more variables simultaneously:

  • Whether a character’s face, clothing, and body shape remain consistent;
  • Whether props appear, disappear, or deform for no reason;
  • Whether a character’s movements flow coherently rather than resetting midway through;
  • Whether camera movement remains continuous and spatial relationships stay stable;
  • Whether visual style, lighting, and color drift during a long take;
  • Whether the events described in the prompt occur in a logical sequence.

This is also why Alibaba is positioning Wan3.0 as a shift from “generating a single shot” to “telling a complete story.” Thirty seconds is already enough for a product feature demonstration, a short instructional explanation, a brief advertising narrative, or a complete sequence consisting of an establishing shot, the main action, and a closing shot.

The model also offers a “smart duration” feature that recommends a more suitable video length based on the prompt. If a single 30-second clip is still not enough, users can continue developing the story with the video extension feature. This design is more sensible than simply requiring users to enter a duration manually. When a prompt describes only a simple action, forcing the model to generate 30 seconds often results in repetition, empty shots, and sluggish pacing, wasting generation resources while making the model’s shortcomings more apparent.

However, 30 seconds is first and foremost a product specification; it does not mean that all 30 seconds will be consistently usable. For video models, the truly valuable metric is not “the maximum number of frames the model can output,” but how many of those frames can go directly onto an editing timeline. Character identity preservation, physical motion, shot continuity, and instruction following still need to be validated through a large number of real-world samples. During the public beta in particular, it is important to distinguish between officially curated examples and ordinary users’ first-generation results.

Document Input Is Even More Worth Watching Than the 30-Second Limit

The update with greater product potential is actually document-to-video generation.

In addition to supporting text, image, audio, and video references, Wan3.0 can read structured content from documents, spreadsheets, and links. The officially listed input formats include:

  • doc, txt, and pdf;
  • xls and numbers;
  • ppt, key, and pages;
  • md;
  • File links or webpage links.

At present, each request can include no more than one file or link. Files must not exceed 100 MB, and documents are limited to 50 pages.

This means users no longer need to manually condense a product manual into a prompt before generating material shot by shot. In theory, they can directly submit a PowerPoint presentation, PDF white paper, or Excel spreadsheet and let the system handle information extraction, script organization, visual presentation, and video generation.

Typical use cases include:

  1. Educational content: Converting course materials into short educational videos with visual demonstrations;
  2. Product introductions: Generating product demos or launch teasers from feature documentation;
  3. Data reporting: Turning trends in spreadsheets into animated charts and narrated visuals;
  4. Corporate training: Converting operating manuals and policy documents into more accessible videos;
  5. Content operations: Rewriting articles or research reports into video formats suitable for content-feed platforms.

This approach bypasses the most crowded area of competition in pure cinematic generation and moves directly into enterprise content production. Enterprises do not lack the occasional stunning AI-generated image; what they truly need is a reliable pipeline that can quickly transform existing documents into videos.

However, it should be made clear that Alibaba has not disclosed the internal technical architecture behind document-to-video generation in this announcement. It may not be as simple as having the video model “natively read” an Excel file. Instead, the process may involve multiple stages, including document parsing, content comprehension, script generation, storyboard planning, asset generation, and audiovisual synthesis. Whether the pipeline uses a single model is not important to end users, but it matters greatly to developers, because an error at any stage could cause the result to deviate from the source document.

For example, year-over-year and period-over-period figures in financial spreadsheets, the hierarchy between titles and notes in PowerPoint presentations, and footnotes and citations in PDFs cannot be glossed over simply by producing something that visually “looks like a business presentation.” Office use cases demand significantly greater factual accuracy than entertainment videos. If Wan3.0 merely turns documents into attractive but inaccurate visuals, its value will be greatly diminished.

Document-to-video generation therefore needs to be evaluated against three key questions:

  • Can it faithfully extract numbers, entities, and logical relationships from documents?
  • Can users edit scripts, storyboards, and data mappings instead of relying on one-shot black-box generation?
  • Can content in the video be traced back to a specific page or spreadsheet cell in the source document?

From an enterprise application perspective, editability and traceability may be even more important than cinematic visuals.

Omni-Reference Aims to Solve the Problem of Characters Changing Between Shots

Wan3.0 has also strengthened its “Omni-Reference” capabilities. According to Alibaba, the model can maintain consistency across characters, props, voices, spatial relationships, and visual styles.

For people, the model attempts to consistently preserve facial features, hairstyles and hair color, body shape, clothing, and accessories. For props, it emphasizes appearance from multiple angles, hardware structure, logos, and material details. For scenes, it must handle the relationship between character positions and camera angles. When multiple styles are involved, it aims to prevent different visual languages from contaminating one another.

These capabilities directly address pain points in commercial video production. In a product advertisement, if the number of cameras on the back of a smartphone changes between angles, or if the brand logo suddenly becomes distorted, the video cannot be delivered no matter how polished it looks. The same applies to character-driven videos: viewers are highly sensitive to changes in faces, and a single obvious instance of identity drift can be enough to break narrative continuity.

For portrait generation, Wan3.0 also proposes “a unique face for every person,” aiming to reduce the plastic-looking skin, formulaic facial features, and exaggerated performances commonly seen in AI-generated characters. The model is designed to handle skin texture, microexpressions, body movements, and the emotions of different people in group scenes with greater nuance.

The direction is sound. Over the past two years, video models have become capable of consistently generating short clips that are impressive at first glance, but their characters still tend to have similarly polished faces and commercial-style expressions. The next stage of differentiation will not be limited to higher resolution. It will depend on whether characters feel specific and authentic, whether their movements are restrained and natural, and whether multiple characters can interact convincingly.

Up to RMB 36 to Generate a 30-Second 1080p Video

Alibaba has announced the API pricing for Wan3.0:

| Output specification | API price | Theoretical cost of generating 30 seconds | | --- | ---: | ---: | | 480p | RMB 0.3/second | RMB 9 | | 720p | RMB 0.6/second | RMB 18 | | 1080p | RMB 1.2/second | RMB 36 |

The “theoretical cost” is calculated based on a single successful 30-second generation. It does not include the cost of generating multiple candidates, retrying failed requests, extending videos, adding voiceovers, or post-production editing.

For developers, per-second pricing is simple and intuitive, but applications cannot consider only the price of a single request. If a finished 1080p video must be generated four times before a usable version can be selected, the generation cost alone rises from RMB 36 to RMB 144. If a product allows users to retry indefinitely, it will also need budget caps, queue priorities, task cancellation, and content moderation mechanisms. Otherwise, usage quotas can be exhausted quickly under high concurrency.

On the other hand, this pricing remains attractive for enterprise short videos and internal training content. Compared with the cost of live-action filming, locations, lighting, actors, and post-production, a generation cost of a few dozen yuan per attempt is not high. For free tools aimed at general consumers, however, continuously providing 30-second 1080p generation still represents a substantial inference cost.

The API is not yet fully available. Details such as invocation methods, concurrency limits, asynchronous task protocols, billing rules for failed requests, generation times, random seeds, content moderation responses, and retention and deletion policies for uploaded documents have yet to be announced. At this stage, it is therefore more appropriate to view Wan3.0 as a preview of its capabilities and a product public beta, rather than as a mature service ready for reliable integration into production environments.

The Real Competition Is Shifting From Image Quality to Workflows

The release of Wan3.0 illustrates a clear change taking place in video generation products: model providers are no longer competing solely over who can generate the most attractive few seconds of footage. They are now competing to own the complete workflow.

The 30-second duration allows the model to support more complete narratives. Smart duration reduces trial and error when configuring parameters. Video extension supports continuous creation. Omni-Reference helps maintain consistency across characters and products. Document input brings an enterprise’s existing materials into the generation workflow. Only when these capabilities are combined does the result begin to resemble a genuinely usable video production system.

The most pragmatic aspect of Alibaba’s approach is that it has not positioned Wan3.0 solely as a filmmaking tool. It has also identified educational courseware, product demonstrations, animated charts, and business presentations as target scenarios. Compared with cinematic long-form video, office videos require less visual spectacle but involve a large volume of clearly defined, repetitive, and monetizable demand, making them more likely to generate significant API usage.

However, document-to-video generation will also take the model into a market with far less tolerance for error. In entertainment content, a cup suddenly changing shape may be only a minor flaw; in a business presentation, misreading a percentage is a factual error. Whether Wan3.0 can progress from being “capable of generation” to being “ready for delivery” will depend on whether it provides supporting features such as script review, shot-level editing, data locking, version management, and source tracking.

This update is therefore worth watching, but it is too early to declare a winner based on specifications alone. The 30-second duration expands the space for expression, while document input expands the sources of content. Reliability, controllability, and the cost per usable finished video will ultimately determine whether Wan3.0 can enter production environments.

Availability and Usage Limits

As of this evening, Wan3.0’s availability is as follows:

  • Alibaba Cloud Model Studio: Public beta available;
  • Wanjing Yike: Public beta available;
  • Official Wanxiang website: Public beta available;
  • Qwen Creation for PC: Public beta available;
  • Qwen app: Phased rollout;
  • API: Pricing announced, with full availability coming soon;
  • Document input: Up to one file or link per request;
  • File limits: No more than 100 MB and no more than 50 pages.

Developers planning to integrate the service can begin by designing their logic for asynchronous generation tasks, asset storage, moderation, retries, and cost controls. However, it is not advisable to commit to a production launch date before the API documentation, service levels, and data processing policies have been fully published.

References

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: