How Google stretched a 10-second model into a 40-second one
The headline "40 seconds" figure is not a single continuous render. Individual generated clips still run 3 to 10 seconds - what changed is the amount of prior video the model can see when extending a clip: up to 10 seconds of context, versus roughly 1 second before [1]. That larger context window lets Omni 1.1 chain clips together in 10-second increments, up to a cumulative 40 seconds [2]. The update also adds first-and-last-frame control, where a user supplies the starting and ending keyframes of a shot and the model fills in continuous motion between them - a feature aimed squarely at camera orbits and other complex movements that are hard to prompt from text alone [1]. On Reddit's r/singularity, the top comment called the wider context window "a giant leap" over models that only reference the last frame, while a separate user who extended an existing Omni 1.0 video reported one specific glitch - an unexpected black glove appearing on a waitress near the end of the clip - a reminder that a bigger context window narrows but doesn't eliminate the model's consistency problems.



