Google Gemini Omni 1.1 Flash Video Model Launch
TECH

Google Gemini Omni 1.1 Flash Video Model Launch

31+
Signals

Strategic Overview

  • 01.
    Google DeepMind announced Gemini Omni 1.1 Flash on August 27, 2026, an updated video generation and editing model available through the Gemini API, Google AI Studio, Google Flow, ComfyUI, and enterprise platforms.
  • 02.
    The model extends generated video in 10-second increments up to a cumulative 40 seconds, analyzing up to 10 seconds of prior context per extension - a major jump from the roughly 1-second context window of the prior version.
  • 03.
    New first-and-last-frame control lets users specify starting and ending keyframes so the model generates continuous video between them, useful for camera orbits and complex movements.
  • 04.
    A new 360p draft mode renders lightweight previews up to 60% faster and at roughly one-third the cost of standard 720p output, intended for rapid iteration before a final render; outputs can then be upscaled to 1080p or 4K.
  • 05.
    Pricing is tiered by resolution: $0.03 per second for 360p, $0.10 for 720p, $0.15 for 1080p, and $0.30 for 4K.
  • 06.
    Every generated video carries Google's SynthID watermark for AI provenance, and Google has proactively restricted the model's ability to alter people's speech during editing despite having the underlying capability.

How Google stretched a 10-second model into a 40-second one

The headline "40 seconds" figure is not a single continuous render. Individual generated clips still run 3 to 10 seconds - what changed is the amount of prior video the model can see when extending a clip: up to 10 seconds of context, versus roughly 1 second before [1]. That larger context window lets Omni 1.1 chain clips together in 10-second increments, up to a cumulative 40 seconds [2]. The update also adds first-and-last-frame control, where a user supplies the starting and ending keyframes of a shot and the model fills in continuous motion between them - a feature aimed squarely at camera orbits and other complex movements that are hard to prompt from text alone [1]. On Reddit's r/singularity, the top comment called the wider context window "a giant leap" over models that only reference the last frame, while a separate user who extended an existing Omni 1.0 video reported one specific glitch - an unexpected black glove appearing on a waitress near the end of the clip - a reminder that a bigger context window narrows but doesn't eliminate the model's consistency problems.

Leaderboard crown, but not a unanimous verdict

Gemini Omni Flash currently leads Artificial Analysis's Text-to-Video Arena (without audio) with an Elo of 1,323 [3], and topped LMArena's Text-to-Video Arena outright with a score of 1,527 [4]. But the picture is less clean on other boards: in the audio-inclusive version of Artificial Analysis's own arena, Omni Flash actually ranks #2 behind Wan 3.0 [3], and in Black Forest Labs' own preliminary human-preference evaluations, its competing FLUX 3 Video model was preferred over Gemini Omni Flash in 52% of head-to-head comparisons [5]. That split matters because it comes from a rival lab's internal testing rather than a neutral benchmark, but it still undercuts any claim that Omni 1.1 Flash is unambiguously the best video model available - the "best" answer depends heavily on which arena, which audio setting, and whose evaluation you trust. The #1 ranking is also the story Google and its own community chose to amplify: the official launch post led with the new controls rather than the benchmark, but independent creators quickly turned it into a head-to-head content genre of its own, pitting Omni 1.1 Flash against rivals like Seedance 2.5 on identical prompts to see whether the leaderboard position holds up to a real side-by-side.

The gap between the spec sheet and what you actually get

Google's own documentation is more careful than its marketing headline suggests: 4K output is explicitly labeled as upscaled, not natively generated - the model renders at lower resolution and then upsamples the result [2]. That same "official spec vs. actual access" gap showed up independently in hands-on testing and community discussion. A YouTube reviewer using the Artlist integration found that platform capped output at 720p and 10-second clips, well below the 40-second, 4K figures in Google's own announcement, suggesting third-party platforms may expose a reduced-spec tier of the model rather than its full capability. Separately, users in Reddit's r/Bard thread reported confusion over whether the new scene-extension feature was even live inside Google Flow at all, with one user unable to find it and another reporting no visible quality improvement in Flow specifically. Three independent sources - a technical breakdown, a hands-on video review, and a user community thread - converging on the same theme (the marketed capability outruns what's actually reachable depending on where you access it) makes this less a one-off complaint and more a real characteristic of the rollout.

Enterprise consolidation, with deliberate guardrails

VentureBeat frames the API release as collapsing what used to be several separate point tools (generation, extension, editing) into a single conversational model, which it argues reduces the number of vendors an enterprise video team needs to manage and simplifies data-handling and output monitoring [4]. Google's own model card tempers that consolidation story with two caveats. First, although the model is technically capable of changing a person's speech during video editing, Google says it is "for now" restricting that capability - an implicit acknowledgment of deepfake and misinformation risk in a tool that otherwise makes video editing conversational and fast [6]. Second, the model card admits that "maintaining complete consistency throughout edits, generating scenes with complex motion, or rendering perfectly accurate text remains a challenge," meaning outputs generated through this consolidated pipeline still need human review before they're production-ready [6].

Historical Context

2026-05-19
The Gemini Omni family, including the base Gemini Omni Flash, was announced at Google I/O 2026 and became generally available the same day via the Gemini app, Google Flow, and YouTube Shorts.
2026-06-30
Gemini Omni Flash was made available to developers for the first time through the Gemini API and Google AI Studio as a preview model.
2026-08-27
Google announced the Gemini Omni 1.1 Flash update, adding longer scene extension, first/last-frame control, 360p draft mode, and 4K upscaling.
prior
Video generation previously required routing through Veo as a separate model and pipeline step; Gemini Omni collapses that into a single natively multimodal model, though Google still positions Veo as a distinct specialized video line.

Power Map

Key Players
Subject

Google Gemini Omni 1.1 Flash Video Model Launch

GO

Google DeepMind

Developer and publisher of Gemini Omni 1.1 Flash; controls model access via the Gemini API, AI Studio, and model card/safety disclosures.

AD

Adobe

Integration partner that has added Gemini Omni Flash into Adobe Firefly for AI video generation.

CO

ComfyUI / Comfy Org

Integration partner shipping partner nodes and workflows that let ComfyUI users call Gemini Omni Flash for text-to-video, image-to-video, and video-editing pipelines.

AR

Artificial Analysis

Independent benchmarking site running blind-vote leaderboards on which Gemini Omni Flash competes against Seedance 2.0, MiniMax H3, and Wan 3.0.

BL

Black Forest Labs

Competitor that ran preliminary human-preference evaluations comparing its FLUX 3 Video model against Gemini Omni Flash.

Fact Check

7 cited
  1. [1] Build with Gemini Omni 1.1 Flash
  2. [2] Google Gemini Omni Flash
  3. [3] Text-to-Video Arena Leaderboard
  4. [4] Google's Gemini Omni Flash Hits the API, Turning Enterprise Video Production Into a Conversation
  5. [5] Gemini Omni Flash Review
  6. [6] Gemini Omni Flash Model Card
  7. [7] Gemini API Docs: Omni

Source Articles

Top 5

THE SIGNAL.

Analysts

Expressed optimism about the model's future trajectory, framing current capability as an early stage of a longer roadmap.

Google Omni team
Google DeepMind (unattributed to a specific individual in the source)

Frames the API release as significant for enterprise video production because it collapses point-tool workflows into a single conversational pipeline, lowering the cost and effort barrier for longer training or explainer videos.

VentureBeat
Technology publication analysis
The Crowd

Gemini Omni 1.1 Flash is our newest multimodal model for video generation and editing. It delivers a new suite of creative capabilities and controls for developers. With this update you can: Extend your scenes, Specify starting and ending frames of a shot, Add video input references, Upscale to 4K, Test ideas quickly in 360p.

@@Google2025

Gemini Omni 1.1 Flash vs Seedance 2.5. Omni 1.1 flash just landed at #1 in the Arena. So... let's put the new champion against Seedance 2.5. Same challenge. Same prompt. Two heavyweights. Can Seedance 2.5 take down the current #1? Watch the comparison and let me know...

@@JSFILMZ0412257

Google just upgraded Gemini Omni Flash with some pretty serious video controls. You can now: extend existing scenes, set the first AND last frame, use reference video for motion + consistency, test ideas cheaply in 360p, upscale the winners to 4K.

@@WesRoth27

Gemini Omni 1.1 Flash now available

@u/PandaElDiablo201
Broadcast
8 Insane Ways to Use Gemini Omni Flash

8 Insane Ways to Use Gemini Omni Flash

I Tested Gemini Omni Flash so You Don't Have to...

I Tested Gemini Omni Flash so You Don't Have to...

NEW AI : Google Gemini Omni Flash | Best AI Video Generator & Editor 2026

NEW AI : Google Gemini Omni Flash | Best AI Video Generator & Editor 2026

Google Gemini Omni 1.1 Flash Video Model Launch — AI News | Agentic Brew