Google releases Gemini Omni 1.1 Flash video generation model
TECH

Google releases Gemini Omni 1.1 Flash video generation model

29+
Signals

Strategic Overview

  • 01.
    Google DeepMind released Gemini Omni 1.1 Flash on August 27, 2026, a production update to its generative video model, available via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.
  • 02.
    Scene extension now analyzes up to 10 seconds of prior context (up from just the final frame previously), extending clips in 10-second increments up to a cumulative 40 seconds, alongside new first-and-last-frame keyframe control.
  • 03.
    A new 360p draft mode runs up to 60% faster and at roughly one-third the cost of standard 720p generation, with outputs upscalable to 1080p or 4K; pricing is tiered from $0.03/sec at 360p up to $0.30/sec at 4K.
  • 04.
    The update also reached Google Flow, giving that consumer/pro video editing product the same scene extension, keyframe control, and 4K export options, available globally to Google AI Plus, Pro, and Ultra subscribers.

Why Scene Extension Finally Works: The Context-Window Fix

The headline technical change in 1.1 Flash is almost boring to describe but hard to overstate in effect: scene extension now reads up to 10 seconds of prior footage before generating the next segment, instead of relying on a single final frame the way the original Omni Flash did [1]. That sounds like a minor parameter tweak, but it directly targets the failure mode that made multi-clip AI video so fragile - a character's jacket changing color, a room's lighting flattening out, or a background prop vanishing between cuts because the model had almost no memory of what came before. With a 10-second window, the model can extend clips in 10-second increments up to a 40-second cumulative maximum, giving creators enough runway to build a short scene rather than a single disconnected shot.

Paired with the context-window upgrade is first-and-last-frame keyframe control, which lets a creator specify exactly how a shot should begin and end - useful for camera orbits, zooms, and looping clips that need to land on a precise final composition rather than wherever the model happens to drift [2]. Google also added support for up to three seconds of reference video, drawn from as many as three separate clips, specifically to help the model hold onto a character's appearance or a camera's motion style across generations [2]. None of this makes Omni cinematic, but it is a meaningful admission that the previous version's memory was the bottleneck, not raw image quality.

That memory upgrade shows up plainly in how creators are actually using the model. Developer walkthroughs posted online have leaned into conversational, iterative editing - asking the model to change a cat's color mid-conversation, for instance, while it keeps the same scene, shot, and camera framing intact. Hands-on demos have also highlighted native audio output generated alongside the video and what creators describe as a kind of world knowledge - motion that respects physics like weight, momentum, and contact rather than sliding or floating unnaturally. A recurring practical use case in these demos is restyling: taking an existing clip and changing its visual style while preserving the underlying character movement, camera path, and composition, which is exactly the kind of workflow the extended context window and reference-video support are built to support.

The Draft-Then-Upscale Economics of Iteration

The more consequential change for anyone actually paying for this model is economic, not creative. A new 360p draft mode generates up to 60% faster and at roughly one-third the cost of standard 720p generation [3], which turns AI video from a one-shot gamble into an iterative workflow: generate cheap, rough drafts at 360p, lock in the composition and motion you want, then pay only once to upscale the winning take to 1080p or 4K. Google's published pricing makes the tiering explicit - $0.03 per second at 360p, $0.10 at 720p, $0.15 at 1080p, and $0.30 at 4K, a 10x spread between the cheapest draft and the most expensive final export [4].

That pricing structure, alongside a separate token-based rate for API calls that mix text, video, and output tokens, is clearly built around a production mentality rather than a one-off novelty generation [1]. For a marketing team or solo creator, it means the cost of exploring five or ten variations on a shot no longer scales linearly with final-render cost - you only pay full price for the version you keep. That draft-first economics is arguably a bigger unlock for adoption than any single creative feature, because it removes the main reason generative video felt too expensive to iterate on.

Google's Unified Multimodal Bet vs Sora 2 and the Asian Video-Gen Wave

Omni Flash's release timing and feature set only make full sense against the competitive backdrop. Commentary framing the model notes a direct contrast with OpenAI's Sora 2: where Sora leans into longer, higher-fidelity premium clips, Omni Flash is being positioned around workflow integration and high-volume, conversational multimodal output rather than chasing raw visual fidelity [6]. That's a different bet - instead of winning on any single generation's polish, Google is trying to win on how many tools a team can retire by using one model for relighting, reframing, wardrobe changes, and text overlays in a single conversational thread [5].

The model's multilingual support (Chinese, Japanese, Korean) is also read as a deliberate move to compete head-on with ByteDance's Seedance, Alibaba's Wan, and Kuaishou's Kling, rather than ceding the fast-growing Asian video-generation market to them [6]. Production integrations reinforce the workflow-first strategy: Adobe has already wired Omni Flash into Firefly for plain-English video edits with provenance metadata carried through [8], and Figma's Weave canvas product cites it as one of the strongest video models it has access to [2]. Both partnerships suggest Google is optimizing for where the model gets embedded, not just how good any single clip looks in isolation.

Benchmark Gold vs the Community Reality Check

There's a real tension between how Omni 1.1 Flash is performing on leaderboards and how it's landing with people actually using it day to day. Third-party benchmarking has been favorable, with the model reported to top text-to-video rankings and place near the top of image-to-video comparisons. But community reaction on forums where creators post side-by-side tests has been decidedly more mixed: the extended context window for scene extension is welcomed as a genuine technical leap that addresses the character- and lighting-consistency problems of the earlier model, yet plenty of hands-on testers still rate the output as noticeably behind Seedance, Minimax, and Kling on physical realism - motion like backflips or fast action still shows artifacts and unnatural slow-motion glitches.

VentureBeat's enterprise-focused analysis lands somewhere in the middle of that divide: it credits the conversational, multi-turn editing as a real productivity win for marketing teams who need quick relighting or wardrobe swaps, while flagging that native clip lengths remain short (three to ten seconds before extension), the base resolution tops out at 720p before upscaling, and there's still no native audio-input support or deepfake/lip-sync capability - meaning premium, longer-form production work is still likely to stay on Veo 3.1 or Sora 2 [5]. The gap between "state of the art on a benchmark" and "as good as the leading Asian video models in practice" is the most honest read on where Omni 1.1 Flash currently sits.

Mandatory Watermarking as the Emerging Provenance Baseline

Every video Omni produces carries Google's imperceptible SynthID digital watermark, with no option in the API to disable it, layered alongside C2PA Content Credentials for provenance tracking [7]. That's a deliberate design choice going back to Omni's original unveiling at Google I/O 2026 by Sundar Pichai and Demis Hassabis, where the model was introduced as Google's first any-to-any multimodal system - and the provenance stack has stayed intact as the model has moved from a consumer-only Gemini app feature into a full developer API and enterprise platform [7]. Notably, the first public hint of the model's existence came weeks earlier, when leak tracker TestingCatalog spotted an unreleased "Powered by Omni" string inside Gemini's video tab [10].

With SynthID and C2PA now baked into every clip with no opt-out, Google is effectively trying to make watermarking a default expectation for AI-generated video rather than an optional add-on - relevant given that OpenAI has also adopted C2PA for its own outputs [7]. As enterprises adopt tools like Omni Flash for real production work, that provenance layer intersects with governance concerns such as the EU AI Act, and analysts have flagged workforce and compliance questions as a medium-term consideration for any company building creative workflows on top of the model [6].

Historical Context

2026-05-02
An unusual UI string reading "Powered by Omni" was spotted inside Gemini's video tab and verified by AI leak tracker TestingCatalog, the first public sign of the Omni model.
2026-05-19
Gemini Omni was officially unveiled at Google I/O 2026 by Sundar Pichai and Demis Hassabis as Google's first any-to-any multimodal model, initially consumer-only via the Gemini app, Google Flow, and free on YouTube Shorts.
2026-06-30
Gemini Omni Flash opened to developers via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.
2026-08-27
Google DeepMind released Gemini Omni 1.1 Flash, adding scene extension to 40 seconds, first/last-frame control, 360p drafting, and 4K upscaling, and bringing the same controls to Google Flow.

Power Map

Key Players
Subject

Google releases Gemini Omni 1.1 Flash video generation model

GO

Google DeepMind

Developer of Gemini Omni 1.1 Flash; positions it as a unified multimodal (text/image/audio/video-in, video-out) model to consolidate video production tooling for enterprise and developer customers.

AD

Adobe (Firefly)

Integrated Gemini Omni Flash into Adobe Firefly, letting editors issue plain-English edit commands with C2PA/SynthID provenance carried through.

FI

Figma (Weave)

Production customer citing Omni Flash as one of the strongest video models available in its Weave canvas product for creative teams.

GO

Google Flow

Google's consumer/pro video editing product that inherited the same scene extension, keyframe, 360p draft, and 4K upscale/export controls as the API model.

OP

OpenAI (Sora 2)

Primary competitor; strategic contrast is visual fidelity/longer premium clips (Sora) vs. workflow integration and high-volume multimodal output (Omni).

BY

ByteDance (Seedance), Alibaba (Wan), Kuaishou (Kling)

Multilingual (Chinese/Japanese/Korean) support signals Google is competing directly against these Asian video-gen platforms rather than ceding that market.

Fact Check

10 cited
  1. [1] Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last-Frame Control, and 4K Upscaling
  2. [2] Build with Gemini Omni 1.1 Flash
  3. [3] Google Flow gets a big Gemini Omni 1.1 Flash video update
  4. [4] Google Gemini Omni 1.1 Flash AI video tools
  5. [5] Google's Gemini Omni Flash hits the API, turning enterprise video production into a conversation
  6. [6] What Gemini Omni signals about Google's AI strategy and the future of multimodal models
  7. [7] Gemini Omni: Google's any-to-any multimodal model
  8. [8] Google Gemini Omni Flash Adobe Firefly video integration
  9. [9] Gemini Omni Flash pricing guide
  10. [10] Google I/O 2026 testing: Omni model spotted early

Source Articles

Top 5

THE SIGNAL.

Analysts

Framed Gemini Omni as a general any-input, any-output creative model starting with video, emphasizing conversational, natural-language video editing.

Koray Kavukcuoglu
CTO of Google DeepMind and Chief AI Architect at Google

Praised Omni Flash's output quality and canvas-integration workflow for creative teams as a production tool.

Itay Schiff
Creative Director, Figma Weave

Argues the model collapses several point tools (relighting, reframing, wardrobe changes) into one conversational workflow, valuable for marketing teams, though 10-second clip length and lack of audio-input/deepfake capability limit premium use.

VentureBeat analysis
Enterprise AI trade press

Interprets Google's unified multimodal architecture as a bet that synchronization quality and workflow simplification (video, voice, music, text in one model) beat marginal per-modality quality gains, targeting enterprise vendor consolidation.

AI Journal analysis
Industry strategy commentary
The Crowd

Rev up your story by extending the scene with Gemini Omni 1.1 Flash. 🏎️ Now you can build longer stories or branch into new creative directions by creating or uploading a video and then asking Gemini to extend the scene from right where it left off.

@@GeminiApp1264

Big news: Gemini Omni 1.1 Flash has landed #1 in the Text-to-Video Arena and #2 in the Image-to-Video Arena! For Text-to-Video the latest 1.1 model is +20pts above FLUX 3 Video at #3 (1495 pts). For Image-to-Video the release is a strong +25pt improvement from Gemini Omni Flash

@@arena1060

Last week, we released Gemini Omni 1.1 Flash, our latest model that gives you more creative controls and generative video capabilities. Omni 1.1 Flash allows for even more creative control, including the ability to extend a scene, first and last frame interpolation, crisp 4K

@@Google369

Gemini Omni 1.1 Flash now available

@u/PandaElDiablo209
Broadcast
8 Insane Ways to Use Gemini Omni Flash

8 Insane Ways to Use Gemini Omni Flash

Google's New Video Model - Gemini Omni Flash - Demo for Devs

Google's New Video Model - Gemini Omni Flash - Demo for Devs

Introducing the Gemini Omni Flash API

Introducing the Gemini Omni Flash API

Google releases Gemini Omni 1.1 Flash video generation model — AI News | Agentic Brew