Google DeepMind
Gemini Omni 1.1 Flash
the production release of Google's video line

Key facts
- 27 Aug 2026generally available
- Released
- #11,514.8 · 30 Aug 2026
- Arena T2V
- #21,487.6 · 30 Aug 2026
- Arena I2V
- 40s10-second extensions
- Cumulative length
- ~$0.10per second of 720p
- Price
- 360pa third of 720p cost
- Draft tier
Google's Omni video model left preview on 27 August 2026. The useful part is control: a first and a last frame, extension that reads ten seconds of what came before rather than one, a 360p draft mode at a third of the price, and 1080p or 4K on the way out. It leads the arena.ai text-to-video board on the 30 August reading, on a small number of votes.
What it is
Gemini Omni 1.1 Flash is the video generation model in Google DeepMind’s Omni line, and on 27 August 2026 Google took it out of preview. The launch post, bylined Anish Nangia and Alisa Fortin, both product managers at Google DeepMind, sets the intent in a sentence: the updates make Omni 1.1 “production-ready for professional use via the Gemini API in Google AI Studio”. Google’s own pricing page carries the same status, describing gemini-omni-1.1-flash as “now generally available to developers on the paid tier of the Gemini API”.
The model takes text, images, video and audio as input and returns video with native audio. Google’s model index gives it one line: “Fast video generation, editing, keyframe interpolation, and extension with native audio”. Read that as the shape of the release. This is a control update rather than a new base model, aimed at developers wiring video into an editing tool or a production pipeline rather than at someone typing a prompt and hoping.
What is actually new
Five capabilities ship together, and four of them are about directing the output rather than improving it.
First and last frame. In Google’s words: “Achieve smooth transitions and camera movements by specifying the starting and ending frames of a shot. Omni 1.1 generates continuous video between two keyframes.” Give it the frame you start on and the frame you land on, and the model fills the middle. For anyone matching a generated shot to plates that already exist, this is the feature that turns a generator into something a cutting room can use.
Scene extension that reads ten seconds back. Earlier models in the line referenced only the final second of a clip when continuing it, which is why extended footage tended to drift away from what it was extending. Omni 1.1 analyses up to ten seconds of prior context, and Google states the limit precisely: “You can extend videos in 10-second increments up to a total cumulative length of 40 seconds.” Forty seconds is not a single-pass generation, and the distinction is worth stating before the comparison below: it is four ten-second passes chained, each one aware of the ten seconds before it.
A 360p draft tier. Google offers “lightweight previews in 360p resolution up to 60% faster* and at a third of the cost compared to Omni 1.1’s standard 720p resolution”, with the asterisk footnoted as a comparison of system throughput at the two resolutions rather than a measured wall-clock result. The workflow this buys is the one professional teams already run in other media: iterate cheap, finish expensive.
Upscaling to 1080p or 4K. The finishing half of that same workflow, run from inside the model rather than through a separate upscaler.
Video references. The capability most write-ups skipped: “Reference up to three seconds of video when crafting your scene.” Style, motion or subject can now be carried in from footage rather than from a still.
The API shape reinforces the pipeline framing. Google’s sample code calls client.interactions.create with the model id gemini-omni-1.1-flash, chains an extension by passing previous_interaction_id, and sets the draft tier through response_format with a resolution of 360p. Drafting and finishing are the same model at a different setting, not two models to integrate.
The scoreboard
On the arena.ai boards, read on 30 August 2026, Gemini Omni 1.1 Flash holds first place on text-to-video with an Elo of 1,514.8. Second is its own predecessor, Gemini Omni Flash, on 1,512.4, and third is Black Forest Labs’ FLUX 3 video model on 1,494.9. On image-to-video it takes second at 1,487.6, behind MiniMax H3 on 1,494.4.
Two numbers deserve as much attention as the ranks. The winning text-to-video row rests on 1,762 votes, while the predecessor sitting 2.4 points below it has 19,830; on image-to-video the new model has 3,720 votes against MiniMax H3’s 27,308. Arena scores settle as votes accumulate, so a fresh model at the top of a board is a real result on a young sample rather than a verdict. Both positions are worth re-reading in a fortnight. Arena publishes no snapshot date of its own, which is why every board figure on this hub carries the day it was taken.
Price and access
There is no free tier. On Google’s published rates, input costs $1.50 per million tokens across text, images, video and audio, and output costs $9.00 per million tokens for text and $17.50 per million for video. Google converts the video rate itself: billing runs at 5,792 tokens per second of 720p video, “an effective price of approximately $0.10 per second” under standard pricing.
That puts it below Google’s own Veo 3.1 at $0.15 a second in fast mode, level with Alibaba’s Wan 3.0 at 720p, and above MiniMax H3’s $0.08 a second at 768p. The 360p draft tier at a third of the 720p cost lands near three cents a second, which changes what a team can afford to throw away.
Access opened on 27 August through Google AI Studio, the Gemini Enterprise Agent Platform API, and Google Flow for AI Plus, Pro and Ultra subscribers worldwide. Scene extension arrived in the Gemini app for the same subscription tiers on the same day. Google names Adobe, which has integrated the model into Firefly, alongside Figma Weave and Runway as early users, and its post carries statements from people at Figma Weave, GMI Cloud and Runway. There are no open weights, and Google claims none.
How it compares
The long-form question separates this release from the Chinese cohort cleanly. ByteDance’s Seedance 2.5 and Alibaba’s Wan 3.0 both generate 30 seconds in a single pass. Gemini Omni 1.1 Flash generates in shorter units and chains them to a cumulative 40 seconds, with each link reading ten seconds of what preceded it. A single pass keeps continuity by construction; a chain keeps it by conditioning. Which suits a job depends on whether the footage needs to be steered part-way through, and the chained approach is the one that lets a director intervene between segments.
On control, the comparison runs the other way. Keyframe interpolation between a supplied first and last frame is not something MiniMax H3 or Seedance 2.5 offer in the same form, and the combination of a cheap draft resolution and an in-model 4K finish is unusual across the field. For a Western buyer weighing procurement, data residency and a billing relationship that already exists, Google remains the low-friction option, and this release closes most of the capability distance that made the Chinese models tempting despite the paperwork.
What to watch
Three things. Whether the arena positions hold once the vote counts on both boards reach the tens of thousands the incumbents carry. Whether Google extends the 40-second cumulative ceiling, which is the one number where the single-pass rivals still read better on a spec sheet. And whether the Omni line now replaces Veo as Google’s default video recommendation, which the pricing and the API surface both point towards without Google having said it outright.
The rest of the field, and how each rival stands on the same boards, is on our AI video models hub, part of our wider AI coverage.
More in Video Generation
All Video →- Google DeepMindVeo 3.1the safest Western default
- AlibabaWan 3.0hosted only, with nothing to download
- MiniMaxMiniMax H3the open flagship
- ByteDanceSeedance 2.0 and 2.5the long-form image-to-video pair
- GuideClaude + MCP for videofrom a written brief to a finished shot
- GuideSeedance 2.5 promptingthe official method, decoded