Technique
First frame and last frame in AI video
How start and end frame video generation works, when to use one frame or two, how to design a pair the model can bridge, and the mistakes that make it morph instead of move.
By Team Pindow · Reviewed
First frame, last frame generation gives a video model two fixed images, where the shot starts and where it must end, and asks it to generate the motion between them. It is the closest thing generated video has to blocking a scene on set: you decide the positions, and the model performs the move.
Used well, it solves three problems at once. It lets a shot land on a precise composition. It lets you control camera moves that text prompts describe unreliably. And it lets you join shots seamlessly, because the end of one clip can be the start of the next. Used badly, it produces morphing: objects melting from one arrangement into another instead of moving. This guide covers how to get the first and avoid the second.
How it works
In ordinary image-to-video generation, you supply a first frame and a prompt, and the model decides where the shot goes. With a last frame as well, the model has a destination. It has to find a plausible path of motion from one image to the other over the length of the clip, guided by your prompt.
That path has to be physically believable for the result to look like movement. If the two frames show the same scene from slightly different positions, or the same character in two poses a few steps apart, the model can bridge them. If they show unrelated compositions, it cannot, and it blends one image into the other instead.
When to use one frame, and when to use two
| Situation | Use | Why |
|---|---|---|
| Ambient motion: wind, water, breathing, a slow drift | First frame only | There is no destination to hit, and a last frame can make natural motion look forced. |
| A character must reach a mark or pose | First and last | The last frame guarantees where they end up. |
| A push in or pull out that ends on a set framing | First and last | Text prompts often overshoot or undershoot a camera move. |
| A reveal, such as a door opening onto a view | First and last | You design the reveal frame instead of hoping for it. |
| Joining two shots without a visible cut | First and last, chained | The end frame of shot A becomes the first frame of shot B. |
| Exploring what a shot could be | First frame only | Let the model surprise you, then pin down the version you want. |
Designing a pair the model can bridge
The quality of the result depends mostly on the two images, not on the prompt. Design them together, from the same reference images and with the same style line.
- Keep the world constant. Same location, same light direction, same time of day, same costume. Only what is supposed to move should differ.
- Make the change a distance a real camera or actor could cover in the clip length. A character crossing half a room in five seconds works. A character appearing on the other side of a city does not.
- Change one thing if you can. A camera move with a still subject, or a subject move with a still camera, bridges far more reliably than both at once.
- Generate the end frame from the start frame. Use the first frame as a reference when you create the last frame, so identity, set dressing and light carry over.
- Watch the edges of frame. Anything that enters or leaves frame between the two images has to be explained by the motion. Unexplained objects are where morphing starts.
What to write in the prompt
With both frames fixed, the prompt only needs to describe how the model should travel between them: the motion, its speed and the camera. Do not redescribe the appearance of the scene. That is already in the images, and restating it can pull the model away from them.
Prompt:
The camera slowly pushes in as she lowers the letter and looks up toward the window. Steady, smooth movement, no cuts. Dust drifts through the light.First frame: a medium shot of her reading at the table. Last frame: a close-up of her face turned toward the window. The prompt names the camera move, the action that connects the two frames and one ambient detail.
- Name the camera move explicitly: "slow push in", "pan left", "static camera". See camera movement for the vocabulary.
- Name the action that connects the frames, in one clause.
- Add "no cuts" or "continuous shot" if the model tends to insert a cut between two different compositions.
- Choose a clip length that suits the distance. Short clips with big changes look rushed; long clips with small changes look sluggish.
Chaining shots into a continuous sequence
To make one shot flow into the next without a cut, take the final frame of the approved take of shot A and use it as the first frame of shot B. Design the last frame of B the same way, and continue. Because each clip starts exactly where the previous one ended, the joins disappear.
Use chaining sparingly. A whole film of invisible joins is tiring to watch and compounds small errors, since each clip inherits any drift in the frame it started from. It is most effective for a single long move, like a camera that follows a character through a doorway and into the next room, or for a hidden cut inside a scene.
Troubleshooting
- The scene morphs instead of moving
- The two frames are too different. Bring them closer together, change only one thing between them, or add an intermediate frame and split the move into two shots.
- The model inserts a cut
- Say "continuous shot, no cuts" in the prompt, and check that the two frames share the same location and light.
- The motion is too fast or jerky
- Increase the clip length, or reduce the distance between the frames.
- The character changes during the move
- Generate the last frame from the first with the character reference attached, and keep the character at a similar size and angle in both frames.
- Nothing happens until the last moment
- Describe the action as starting immediately ("as the shot begins, she turns") so the model spreads the motion across the clip.
Which models support it
Support varies by model and changes as providers update them. At the time of review, several video models on Pindow accept both a start and an end frame, including Veo 3.1, Kling 3.0, Seedance 2.0 and LTX-2.3. Some go further and accept several keyframes spread across the clip, which lets you pin more than two moments. Others take a start frame only. Check the settings of the model you plan to use before designing a shot around a last frame, and see the models page for the current list.
First and last frames work best when the frames come from a proper storyboard. If you have not made one, start with how to storyboard a film with AI, and keep your characters steady across both frames with the consistent characters method.
Try it on Pindow
Pindow is in early access. Request an invite and you can run every step in this guide on one canvas.
Keep reading
- Pre-productionHow to storyboard a film with AI
- TechniqueHow to keep characters consistent in AI video
- WorkflowThe AI filmmaking workflow, step by step
- GuideAI filmmaking: a practical guide from idea to finished film
Back to Learn AI filmmaking

