←Back to News
RESEARCH—Aug 27, 2026—9 Min Read

Ten Seconds of Memory

Model Gemini Omni 1.1 Flash · Launched 27 August 2026 · Team Pindow · 9 min

Pindow and Gemini Omni

Omni 1.1 Flash doesn’t make longer videos. It makes videos that remember what just happened, which turns out to be a different and better problem to solve.

There is a particular failure that anyone who has extended an AI clip knows by heart. You have a good ten seconds. You ask for ten more. And somewhere across the join, the world quietly changes its mind. The wall behind her shifts a shade cooler. The bokeh in the background resolves into different shapes. The hand that was resting on the table is now, inexplicably, not.

The cause was never mysterious. Extension models looked at the last frame and continued from there. One frame is not a scene. One frame has no idea that the light is coming from a window camera left, that the actor has been walking at a constant speed, or that the sound of traffic has been building for six seconds. It only knows what the final twenty-fourth of a second looked like.

Google’s 27 August release changes that number from one frame to ten seconds. Everything interesting about this model follows from that single decision.

What actually shipped

Gemini Omni 1.1 Flash went generally available on 27 August 2026 through the Gemini API, Google AI Studio and the Gemini Enterprise Agent Platform, with the features surfacing in Google Flow for Plus, Pro and Ultra subscribers and scene extension arriving in the Gemini app.

The capabilities:

  • Base generation of 3 to 10 second clips at 24 fps. Not 30, not 25. Worth noting that the default is the cinema rate.
  • Extension in 10-second increments, to a 40-second total. Critically, the model analyses up to 10 seconds of prior context when extending, where earlier versions referenced only the final second.
  • First-and-last-frame control. Specify both ends of a shot as images and let the model generate the transition between them.
  • Video reference input of up to 3 seconds, to carry a character or a motion forward.
  • Resolutions at 360p, 720p native, and 1080p and 4K upscaled.
  • A draft mode at 360p, running up to 60 percent faster.
  • Audio generated alongside the video.

One housekeeping item that will bite somebody: the preview endpoint is deprecated on 30 September 2026. If you built against preview, you have days rather than weeks. Also note that uploaded-video editing is not supported in every region, so check before you promise a client a turnaround.

What it changes on the floor

Forty seconds is a scene, not a shot. Ten seconds is a shot. Twenty is a shot that is overstaying. Forty seconds, held together with consistent light and geography, is enough for an entrance, a beat and a turn, which is to say enough for drama. The unit of AI video has quietly moved up a level.

Ten seconds of context is the whole ballgame. This is the least glamorous line in Google’s announcement and the most important one. A model that reads ten seconds of history can infer velocity, light direction, the rhythm of a movement, the trajectory of a camera move. It can continue a dolly at the same speed. It can keep the sun where it was. One-frame extension could never do any of that, because none of that information exists in a single frame.

First-and-last-frame control is the storyboard, made executable. Every director already works this way. You know the frame you are starting on and the frame you are landing on, and the shot is the journey between them. Handing the model both ends and letting it solve the middle is a far more natural instruction than describing a camera move in prose and hoping. It is also, straightforwardly, how a previz artist thinks.

Draft mode restores the rehearsal. A fast, light 360p pass is not a compromised output. It is a rehearsal take. Block the shot. Find out whether the geometry works, whether the move lands, whether the timing breathes. Then roll the real thing. Every production discipline worth having came from the fact that film used to be consumed by the foot. Draft modes give that discipline back to a process that had lost it.

The resolution ladder is unusually legible. 360p through to 4K with clean steps between, which makes planning a sequence something you can do on paper before you start. That is rarer than it should be.

The grammar test

For extension models we run the continuous dolly.

Generate a 10-second shot with a slow, constant-speed push-in on a subject, with a clearly motivated hard light source on one side of the frame and a practical visible in the background. Then extend it twice, to 30 seconds. Change nothing in the prompt except the instruction to continue.

Watch four things across the joins.

  1. Does the dolly hold its speed? Most models decelerate at the join and then re-accelerate, producing a lurch you cannot grade out.
  2. Does the key light stay on the same side of the face? This is the one that ruins cuts. A light that migrates across an extend is a light that will not intercut with anything.
  3. Does the background practical stay put in three-dimensional space as the camera moves, or does it slide in the frame as though painted on glass?
  4. Does the audio bed continue, or does it restart?

Omni 1.1 Flash passes one and two consistently in our runs. The speed holds and the key stays honest, which is what ten seconds of context buys you. Three is much improved and still occasionally reveals that the model is solving a two-dimensional problem. The parallax on background elements is plausible rather than correct, and on a long push you can catch it. Four is inconsistent enough that we treat generated audio as a temp track rather than a stem.

That third point is the honest ceiling of every video model in this series. They are extraordinarily good at what a move looks like and still approximate about what a space is. This is precisely why a 3D blocking stage exists as its own step in a serious pipeline rather than as a prompt adjective.

Where it breaks

  • Parallax is plausible rather than measured. Long moves through deep space will betray it. Short moves and shallow spaces are fine.
  • 40 seconds is a ceiling rather than a target. Quality degrades toward the end of the chain. Two extends are comfortable. Three is where we start seeing the world drift.
  • Upscaled is not native. 1080p and 4K are upscales from the 720p native tier. They are good upscales. They are not the same as native capture, and on fine texture like hair, fabric weave and foliage you will see the difference on a large screen.
  • Audio is a temp track. Usable for timing. Not usable for delivery.
  • The 30 September preview deprecation will break somebody’s pipeline. Do not let it be yours.
  • Regional limits on uploaded-video editing. Verify availability before quoting a client.

Inside Pindow

Straight into Canvas, and specifically into the relationship between Canvas and the 3D Stage.

First-and-last-frame control is the feature that makes the 3D Stage worth its existence to people who previously thought of it as an extra step. Block the shot spatially: where the camera starts, where it ends, where the subject is, where the light is. Export the two ends as frames. Hand the model a start and a finish that are geometrically correct rather than described, and the middle it invents is constrained by real spatial truth instead of by an adjective.

That is the entire argument for having a blocking stage at all, and it is an old argument. Directors have been drawing the first and last frame of a shot on the back of a call sheet for a hundred years. The tool is new. The habit is not.

On the node graph, the practical arrangement is 360p draft nodes for blocking, 720p for the approved take, and upscaling at the end of the chain rather than in the middle. Upscaling mid-graph and then extending from the upscale gives you the worst of both, with the artefacts of the upscale baked into everything downstream.

Working notes

Extend with the same prompt rather than a new one. The temptation is to describe what should happen next. Resist it. The model’s ten seconds of context is a stronger signal than your sentence, and a new prompt fights the continuity you are trying to buy.

Put the light source in the prompt every single time. Colour temperature, direction, quality. If you do not specify it, the model will re-decide, and a light that re-decides across an extend is a shot you cannot use.

Use first-and-last-frame for anything with a camera move. Prose descriptions of camera moves are the least reliable instruction in AI filmmaking. Two frames are unambiguous.

Draft everything at 360p. Everything. There is no shot so simple that a rehearsal is wasted.

Treat the 3-second video reference as a motion reference rather than a style reference. It carries movement and character well. It carries grade poorly. Do your grade downstream where you can see it.

Ten seconds of memory is a modest-sounding upgrade and it is, in fact, the whole thing. A camera that remembers where it was is a camera you can follow with. That is not a feature. It is the minimum requirement for a shot to become a scene.

Just make a film.

Sources

Gemini Omni 1.1 Flash lets you build with more control (Google) · Gemini Omni 1.1 Flash Preview (Google Cloud documentation) · Google Upgrades Gemini AI Video: What Omni 1.1 Changes for Creators (TechRepublic) · Google Launches Gemini Omni 1.1 Flash With Greater Video Creation Controls (ETV Bharat)

Related Stories

View all →