Direct Veo 3.1 Like a Shot List: A Practical Guide to First-and-Last-Frame Video First-and-last-f...Direct Veo 3.1 Like a Shot List: A Practical Guide to First-and-Last-Frame Video First-and-last-f...
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
Direct Veo 3.1 Like a Shot List: A Practical Guide to First-and-Last-Frame Video
First-and-last-frame control changes the way an AI video can be planned. Instead of describing an entire clip and accepting wherever it ends, you can define the opening composition, define the closing composition, and ask the model to create the transition between them.
That control is powerful, but it does not remove the need for direction. The two images must describe a believable visual journey, and the prompt must explain how the scene moves through time.
Begin with the Story Beat
Before choosing images, summarize the shot in one sentence.
“A closed package opens to reveal the product.”
“A daytime street becomes the same street at night.”
“A wide view of a character ends on a close emotional reaction.”
“A rough sketch becomes a finished design.”
This sentence is the story beat. It helps determine whether the shot should use text-to-video, a single image reference, or first-and-last-frame control.
Use first-and-last frames when the ending matters. If only the general style or subject matters, a standard text or image-to-video workflow may be simpler.
Design the First Frame as a Beginning
The first frame should contain a stable composition that can naturally initiate movement. Leave space in the direction of travel. Make sure important subjects are visible, separated from the background, and large enough to read.
A first frame is not merely an attractive image. It is the opening state of an action.
For a product reveal, show the unopened package clearly and leave visual room for the lid, drawer, or camera to move. For a character shot, use a pose that can transition into the intended final pose without requiring the body to completely reorganize.
Avoid beginning with cropped hands, hidden product edges, or extreme perspective unless those choices are essential to the shot.
Design the Last Frame as a Resolution
The last frame should feel like a plausible consequence of the first. It can change position, camera distance, lighting, environment, or subject state, but it should preserve enough visual continuity for the model to connect the two images.
Strong pairs share several anchors:
The same subject identity.
Compatible proportions.
A related camera axis.
A recognizable environment or color palette.
A clear direction of change.
The last frame should also work as an ending. A product may settle into a clean hero angle. A character may complete a turn and hold the expression. A landscape may reach its final lighting state.
The goal is not to make both frames nearly identical. It is to make the transformation legible.
Control the Size of the Visual Gap
The greater the difference between the first and last frame, the more the model must invent.
A small gap might change a facial expression, move the camera closer, or shift from daylight to golden hour. A medium gap might open a package, rotate a product, or move a character across the frame. A large gap might transform one object into another, move to a completely different location, and change the camera angle at the same time.
Large gaps can produce exciting results, but they are less predictable. If continuity matters, divide one ambitious transformation into multiple short shots.
For example:
Shot 1: Closed package to partially opened package.
Shot 2: Partially opened package to full product reveal.
Shot 3: Product reveal to final logo composition.
Each transition now has a clearer visual task.
Keep Identity Anchors Consistent
When the same person, object, or brand appears in both frames, compare the identity-defining details before generation.
For a person: face shape, hairstyle, age, clothing, accessories, and body proportions.
For a product: silhouette, logo, label, color, material, and small functional details.
For a designed character: line style, costume, palette, and distinctive features.
Differences between the images may turn into morphing during the transition. If a label moves, a hairstyle changes, or a sleeve switches color, the video has to reconcile that inconsistency.
Use reference images to reinforce style and appearance when the workflow supports them. Veo 3.1 can use multi-image reference input to guide visual style, color palette, and composition, but references should agree with one another rather than introduce competing designs.
Write the Prompt for the Journey
The images define where the shot begins and ends. The prompt should describe what happens between those points.
A useful structure is:
Preservation: What must remain consistent?
Action: What changes?
Camera: How does the viewer move?
Timing: How is the action paced?
Environment: What secondary motion supports it?
Audio: What should be heard?
Ending: How does the shot settle?
Example:
“Preserve the exact product shape, label, and deep green color. The box opens slowly as the camera makes a smooth dolly-in. Soft warm light expands from inside the package, revealing the bottle without changing its design. Keep the movement controlled and premium. Add a subtle cardboard opening sound and a low ambient tone. End on the supplied hero frame with the product sharp and centered.”
The prompt does not need to redescribe every pixel in the start and end images. It needs to explain the transition.
Direct Timing Inside the Clip
Veo 3.1 produces short clips, so every instruction competes for time. A shot with five separate events may rush through each one.
Think in simple beats:
Opening: Establish the first frame.
Middle: Perform the main change.
Closing: Arrive at the final frame and hold briefly.
For an eight-second clip, a practical rhythm might be two seconds of establishment, four seconds of transition, and two seconds of resolution. This is a creative guide rather than a frame-accurate command, but it helps keep the prompt focused.
Words such as slowly, gradually, immediately, then, as, and finally can clarify order. Avoid long lists of simultaneous actions unless the visual chaos is intentional.
Use Camera Direction to Connect the Frames
The camera should explain the compositional difference between the two images.
If the last frame is closer, use a dolly-in or controlled zoom. If the subject moves from left to right, use a tracking movement. If the final frame reveals height, use a tilt or crane. If the two frames show different sides of a product, use a restrained orbit.
Do not request a camera move that conflicts with the supplied endpoints. A prompt asking for a leftward pan while the final composition reveals space on the right creates contradictory guidance.
Strong camera instructions are visible and singular: slow dolly-in, lateral track, gentle orbit, static locked shot, handheld drift, or crane up.
Plan Native Audio as Part of the Scene
Veo 3.1 can generate synchronized audio, including dialogue, ambient sound, and effects. Audio should support the same story beat as the image transition.
Separate audio layers:
Primary sound: the action itself, such as a door closing, footsteps, fabric movement, or a package opening.
Ambient sound: rain, room tone, wind, traffic, or a crowd.
Dialogue: exact words, speaker, tone, and timing.
Music: only when it helps the shot, described by mood and intensity.
A concise direction is more useful than an overloaded sound brief: “Soft rain ambience, one clear footstep sequence, and a quiet door latch at the end. No music.”
If dialogue is important, keep it short enough for the clip and avoid asking for several speakers to overlap.
Choose Aspect Ratio Before Building the Frames
Veo 3.1 supports landscape and vertical workflows. A 16:9 composition is suited to wide environments, presentations, and YouTube. A 9:16 composition is suited to Shorts, Reels, TikTok, and subject-centered vertical movement.
Create both endpoint images in the intended ratio. Cropping them later can remove the motion path or change the perceived camera move.
For vertical video, protect space near the top and bottom for platform UI and captions. For landscape video, use lateral space intentionally rather than leaving the subject floating in the center.
Use Fast Iteration to Test the Transition
Do not treat the first generation as the final render. Test the relationship between the two frames before polishing the prompt.
Review:
Does the subject remain recognizable?
Does the movement travel in the intended direction?
Does the model arrive at the last frame naturally?
Is the transition too busy or too slow?
Does the audio support the visual action?
Are logos, faces, hands, and small details stable?
Change one variable at a time. Simplify the action before adding more style language. If the visual gap is too large, create an intermediate frame or divide the idea into separate clips.
A browser-based Veo 3.1 workspace such as VOE31 provides text-to-video, first-and-last-frame, and image-reference modes in one interface. It also exposes native audio and 1080p output options, making it possible to test the story structure before moving the result into a longer editing timeline.
Build Longer Sequences from Short Shots
A coherent short shot can become a building block for a longer video. Plan each clip around one action and make the ending composition useful as the beginning of the next.
A simple sequence might be:
Establish the location.
Introduce the subject.
Perform the main action.
Reveal the result.
Finish on a clean brand or narrative image.
Reuse visual anchors across clips: palette, lighting direction, lens feeling, character design, wardrobe, and sound atmosphere. Continuity is easier to maintain when those decisions are written into a small style guide before generation.
Know the Limits of Endpoint Control
First-and-last-frame control guides a transition; it does not create a frame-perfect animation curve. Hidden product surfaces, unseen parts of a costume, and complex physical interactions still require interpretation.
The technique is well suited to concept videos, ad transitions, product reveals, time-of-day changes, morphs, storyboards, and previsualization. Work requiring exact legal packaging, technical demonstrations, or identity-critical continuity needs careful human review.
A First-and-Last-Frame Checklist
Before generating, confirm:
The shot has one clear story beat.
Both frames use the intended aspect ratio.
The same subject remains recognizable in both images.
Logos, clothing, colors, and proportions are consistent.
The camera axis and perspective can be connected plausibly.
The visual gap fits within one short clip.
The prompt describes the journey rather than repeating the images.
There is one primary camera move.
Action timing includes an opening, transition, and resolution.
Audio instructions are short and compatible with the scene.
The final frame works as a deliberate ending.
All source and reference images are authorized for use.
Endpoints Turn Generation into Direction
Text prompting asks a model to imagine a shot. First-and-last-frame control gives that shot a destination.
The most reliable workflow treats the two frames as production design and the prompt as direction. Choose a believable beginning, a deliberate ending, and one clear visual journey between them. That structure gives Veo 3.1 a stronger foundation for motion, continuity, camera behavior, and sound.

voe31.com

Free Veo 3.1 AI Video Generator — First & Last Frame Control

Create cinematic 1080p videos with Google DeepMind’s Veo 3.1 — text-to-video, first & last frame control, native audio. Sign up free — trial credits included.

Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started