Alibaba
Wan 3.0
hosted only, with nothing to download

Key facts
- 24 Aug 2026public beta 6 August
- Released
- #11,414.1 · 30 Aug 2026
- Arena editing
- 30ssingle pass, 30fps
- Max length
- Nonehosted access only
- Weights
- $0.20per second
- 1080P price
- PPT · PDF · XLS1 file, 100MB, 50 pages
- Document input
Alibaba's newest video flagship generates 30 seconds in one pass, takes a slide deck or a spreadsheet as the brief, and tops the arena.ai video-editing board on the 30 August 2026 reading. It also carries on the family's newer habit: there are no Wan 3.0 weights to download, and the openly published Wan line still stops at 2.2.
What it is
Wan 3.0, which Alibaba writes closed up as Wan3.0, is the current flagship of the Wan video family. Alibaba Cloud announced a public beta on 6 August 2026 and the Wan account declared it generally available on 24 August, with a launch discount of 30 per cent on the standard tier running from 23 August to 23 September 2026 across Alibaba Cloud Model Studio and Qwen Cloud. Alibaba counts eight iterations from Wan 1.0 to here.
The framing Alibaba gives it is broader than text-to-video. The model generates up to 30 seconds in a single pass at 30 frames per second, and it accepts text, images, audio and video as input, plus a category none of its rivals take: documents. A slide deck, a PDF, a word-processed report or a spreadsheet can be the brief. Alibaba’s own examples of what that is for read as a list of jobs rather than a demo: “A product PPT becomes a brand film or product launch ad. A training deck becomes video courseware. Spreadsheet data becomes animated, dynamic charts. A business report becomes a narrated video briefing.”
Where the open line stopped
There are no Wan 3.0 weights, and this release did not start that. Alibaba’s blog makes no open-source claim in its text or its FAQ, and the Wan organisation on Hugging Face has published nothing newer than the Wan 2.2 line: the most recent uploads there are Wan2.2-Animate-2-14B in its Diffusers and Distilled variants, both last modified on 13 August 2026, with Wan-Dancer-14B from 17 July. Checked on 30 August 2026, the openly downloadable Wan weights still stop at 2.2, under Apache 2.0.
The pattern was set before this release. Alibaba published Wan 2.1 and Wan 2.2 openly and then stopped, and our page on Wan 2.7 records the same arrangement a generation earlier: a flagship reachable through hosted APIs while the downloadable weights sat several versions behind it. Wan 3.0 continues that, rather than introducing it. What still trades on the older releases is the family’s reputation, built on permissive licensing and a compact variant that runs on a consumer graphics card, and that reputation is why people arrive at a Wan page expecting a download. The download is 2.2 and older.
So the practical position at 3.0 is the one Wan 2.7 already set. This is a hosted API, priced per second, reachable only through Alibaba Cloud, and anyone who chose the family for self-hosting is not being upgraded by it. The self-hosting advice on the Wan 2.7 page is unchanged for exactly this reason.
Whether the weights follow later is an open question. Alibaba has said nothing either way about 3.0, and it has published openly before, so a delayed release is possible. A page like this can only report the position as it stands, and as it stands there is nothing to download.
The numbers, as Alibaba documents them
The Model Studio API reference, last updated 28 August 2026, is where the specifications live. Two model ids are documented: wan3.0-video as the standard model, and wan3.0-video-prime, which Alibaba describes as a “High-speed version with capabilities aligned to the standard version, with significantly improved end-to-end speed.”
Duration is a parameter taking any value from 2 to 30 seconds when no video is supplied, or -1 for smart duration, in which case the model reads the prompt and proposes a length itself. Output resolutions are 480P, 720P and 1080P, with 1080P the default. Input media types are first_frame, last_frame, reference_image, reference_video, reference_audio, file and link, alongside the text prompt. Document formats accepted are doc, xls, ppt, pdf, txt, key, pages, numbers and md, limited to one file or link at a time, 100MB and 50 pages.
Pricing is by the second and by resolution: $0.05 at 480P, $0.10 at 720P and $0.20 at 1080P. A 30-second generation at 1080P therefore costs $6.00, and the same 30 seconds at 480P costs $1.50, which is the draft-then-finish pattern the tiers are built for.
Alibaba publishes no benchmark table for the model, and it states its own weak points: “Audio texture and on-screen text rendering accuracy are still improving.”
The scoreboard
On the arena.ai boards, read on 30 August 2026, Wan 3.0 holds first place on video editing with an Elo of 1,414.1, ahead of ByteDance’s Seedance 2.5 at 1,409.8. The margin is 4.3 points on 463 votes for the leader and 429 for the runner-up, which are small samples by the standards of that site: read the lead as provisional and re-check it. The ranking is our measurement rather than Alibaba’s claim, and arena publishes no snapshot date, so every board figure on this hub carries the day it was taken.
Video editing is the right board for it. Editing came into the family with Wan 2.7, which allowed visuals, plot and dialogue to be changed without regenerating from scratch, and 3.0 carries that forward. Taking the top slot on the board that measures precisely that capability is a coherent result rather than a surprise.
What changed against Wan 2.7
Alibaba’s own answer runs to four points. Single-generation length extends to 30 seconds, supported by smart duration and by a video-extension path that continues a storyline past the first generation. Document input arrives, which is the feature nothing else in the field offers. Realism is upgraded, with the company citing more varied and lifelike faces and finer-grained reference consistency across characters, props, spaces and style. And the video editing introduced in Wan 2.7 carries over intact.
The two length features work together. Thirty seconds in a single pass is long enough for a complete short-form piece, and smart duration means a user who has not thought about length gets one chosen from the prompt rather than a default that truncates the idea.
How it sits against the field
Against Google’s Gemini Omni 1.1 Flash, released three days later, the two models solve the long-form problem differently: Wan 3.0 generates 30 seconds in one pass, while the Google model chains 10-second extensions to a cumulative 40 seconds. At 720p the pricing is level, both at about $0.10 a second. Against MiniMax H3, the contrast is sharper: H3 publishes its weights, subject to a licence that excludes the United Kingdom, the European Union, the United States and South Korea, and generates 4 to 15 second clips. Wan 3.0 publishes nothing and generates twice the length.
For a UK team, the practical position is straightforward. Wan 3.0 is an Alibaba Cloud service, so the questions to answer are the familiar ones about data residency and procurement policy rather than about licences, and the document input is a genuine reason to try it: turning an approved deck into a video briefing is a task most organisations have and few tools do well. For a team that came to Wan because the weights were free, the answer is that the free path stops at 2.2.
What to watch
Whether Alibaba publishes Wan 3.0 weights at any point, which would restore the family’s position at the open end of this section. Whether the video-editing lead survives a larger vote sample. And whether the document path turns out to be a serious workflow or a demonstration, which will show up in how quickly the on-screen text rendering Alibaba flagged as still improving actually improves, because a slide deck turned into video is mostly text on screen.
The full field, and how each rival stands on the same boards, is on our AI video models hub, part of our wider AI coverage.
More in Video Generation
All Video →- AlibabaWan 2.7the open workhorse
- ByteDanceSeedance 2.0 and 2.5the long-form image-to-video pair
- MiniMaxMiniMax H3the open flagship
- Google DeepMindGemini Omni 1.1 Flashthe production release of Google's video line
- GuideClaude + MCP for videofrom a written brief to a finished shot
- GuideSeedance 2.5 promptingthe official method, decoded