Alibaba's Wan 3.0 generates 30-second videos from prompts and documents

The public beta doubles Wan's prior 15-second ceiling and starts at 5 cents per generated second through Alibaba Cloud and Qwen Cloud.

By · Published

Why it matters

Wan 3.0 doubles Alibaba's prior generation limit while turning decks, spreadsheets and webpages into video inputs, pushing the product closer to automated commercial production.

Alibaba's Wan 3.0 generates 30-second videos from prompts and documents — The public beta doubles Wan's prior 15-second ceiling and starts at 5 cents per generated second through Alibaba Cloud and Qwen Cloud.

Alibaba's Wan team released Wan 3.0 in public beta on August 6th, extending its AI video generator to 30-second clips and letting users build videos from documents, spreadsheets, presentations and webpages alongside conventional text and media prompts.

The model is available through Alibaba Cloud Model Studio, Qwen Cloud and the Wan creation site. The launch makes Wan 3.0 a hosted product first, giving Alibaba a direct route to charge for generation while it tests demand and capacity under public-beta limits.

Wan 3.0 costs $0.05 per generated second at 480p, $0.10 at 720p and $0.20 at 1080p, according to Qwen Cloud's model page. A full 30-second generation therefore costs $1.50, $3 or $6, respectively, before a user pays for retries or alternate takes. Qwen Cloud currently lists a limit of two concurrent jobs, 30 requests per minute and 50 tasks in the asynchronous queue. (qwencloud.com)

Alibaba doubles Wan's single-run duration

Duration is the release's clearest technical and commercial change. Alibaba's previous Wan 2.7 text-to-video and image-to-video models generated clips lasting between two and 15 seconds. Wan 3.0 raises that ceiling to 30 seconds in one run, reducing the need to join separately generated segments that may drift in character appearance, lighting or scene composition. (docs.qwencloud.com)

The longer format also changes the economics of experimentation. A developer could generate a 15-second Wan 2.7 clip at 720p for $1.50. Wan 3.0 preserves the same 10-cent per-second rate at that resolution, so Alibaba is charging linearly for the added duration rather than attaching a premium to the longer context window.

Alibaba describes Wan 3.0 as an all-in-one model for generation, editing, visual replication and character driving. It accepts text, images, audio and video, while its new "Omni-Reference" feature can parse PDFs, text files, Microsoft Office documents, Apple Keynote and Pages files, and webpages. The model's Qwen Cloud listing says it can use those materials as references while maintaining characters and producing synchronized visuals and sound. (qwencloud.com)

The document support is the more consequential product bet. A marketing team could feed Wan 3.0 a product page or presentation instead of manually converting the material into prompts and reference images. A company could use a campaign brief, sales deck or spreadsheet as source material for an initial video. Those workflows depend on how accurately the model extracts facts, layouts and visual priorities from structured files, an area that cannot be judged from Alibaba's launch demos alone.

Alibaba also promotes "reality-grade rendering," including more expressive characters, stronger reference consistency and improved rendering of digital interfaces. That language remains Alibaba's characterization of the output. The measurable parts of the launch are the 30-second duration, expanded reference formats, hosted availability and published API price.

Wan moves from open models to a metered production service

Wan emerged from Alibaba's Tongyi Lab as its visual-generation model family. The team built an early developer following through downloadable releases: Alibaba open-sourced Wan 2.1 in February 2025 and Wan 2.2 that July, later releasing specialized speech-to-video and character-animation variants. Alibaba said in August 2025 that the Wan series had accumulated 6.9 million downloads across Hugging Face and ModelScope. (alibabacloud.com)

That history gave developers access to model weights and local inference code, including a 5-billion-parameter Wan 2.2 model and larger mixture-of-experts variants. Wan 3.0's public-beta rollout instead directs users to Alibaba's own interfaces and API endpoints, where every generated second produces cloud revenue and Alibaba controls the available compute.

The API-first approach also gives Alibaba usage data from document-driven and longer-form generation, the two features that distinguish Wan 3.0 from its immediate predecessor. The company can see which input formats customers use, how often 30-second jobs require retries and whether users accept a $6 price for a full-length 1080p result.

Wan 3.0 arrives as video labs compete to turn visually impressive demonstrations into repeatable production tools. Alibaba is placing its bet on fewer stitched clips, broader source material and transparent per-second pricing. The public beta will test whether those controls can make generated video dependable enough for developers and creative teams to build into routine workflows.

Reader comments

Conversation for this story loads after sign-in.