AI Film ContestsβΊGuidesβΊHow to Make a Clean First Frame for Image-to-Video
How to Make a Clean First Frame for Image-to-Video
In image-to-video, the still you upload becomes the start of the shot, so every choice in it carries into every frame: the face, the product, the light, the palette and the composition.
A clean first frame matches the output's aspect ratio and resolution, has no text or logos baked in, is sharp, leaves room for the motion, and already matches the film's look. Then the text prompt only has to describe what moves.
This guide collects each major tool's rules for input images as of September 24, 2026, a ten-point check for the frame itself, what the vendors say to put in the motion prompt, and when a last frame helps.
What the first frame decides
A video model reads the image as the answer to most of your questions. It keeps the composition, the light, the colors, the wardrobe and the product, and it invents the motion and anything the camera reveals. Every flaw in the still, a sixth finger, a warped label, a face that is slightly off, becomes a moving flaw in every frame after it. Fixing a still costs cents. Fixing a clip means paying for the clip again.
That is why image-to-video gives you so much control over an AI film or commercial. You approve the shot as a still, where problems are cheap to see and fix, and spend on video only for frames you already like.
What each tool accepts as a first frame
The rules differ more than you would expect, so check them before you size the frame. As of September 24, 2026:
- Runway Gen-4.5 accepts JPEG, PNG and WebP but not GIF, in aspect ratios from 1:2 to 2:1, and crops anything else from the center. "Your input image acts as the first frame." Clips run 2 to 10 seconds.
- Kling 3.0 takes JPG or PNG up to 50 MB, at least 300 pixels on a side, from 1:2.5 to 2.5:1, for clips of 3 to 15 seconds up to 4K. It supports start and end frames, but not an end frame on its own.
- Google Veo 3.1 takes JPEG or PNG up to 20 MB and "uses the input image as the initial frame," with an optional last frame. Other sizes "may be resized or centrally cropped," and Google advises 720p or higher at 16:9 or 9:16. Clips are 4, 6 or 8 seconds, and 1080p and 4K come only at 8. Gemini Omni Flash, now Google's default video model in the Gemini API, decides how to use an image unless you tag it as the first frame.
- Seedance on BytePlus ModelArk accepts JPEG, PNG, WebP, BMP, TIFF and GIF under 30 MB, 300 to 6,000 pixels on a side, in ratios between 0.4 and 2.5, and center-crops a mismatch. Seedance 2.0 makes 4 to 15 second clips and Seedance 2.5 up to 30.
- Luma Ray3.2 takes start and end frames in JPEG, PNG or WebP up to 50 MB and 8,000 pixels, for 5 or 10 second clips up to 1080p; start and end frames do not work with 10-second clips.
- MiniMax's Hailuo API takes one first frame, one last frame and up to nine reference images, each up to 30 MB.
- OpenAI removed Sora 2 and its Videos API from the API on September 24, 2026.
The vendors also name two failure modes. Runway's guide warns that in image-to-video, "Visual artifacts, such as blurry hands or faces, may be intensified." ByteDance warns that when input and output sizes do not match, Seedance 2.0 can show "abrupt changes such as image stretching and compression." Both are problems you fix in the still, before you pay for the clip.
Ten checks before you upload
- Match the aspect ratio of the output. Crop the still yourself to 16:9, 9:16 or whatever you will deliver, so the tool never crops or pads it for you.
- Match or beat the output resolution. A frame smaller than the video gives the model less detail to keep.
- Keep text, logos, captions and watermarks out of the frame. Small type warps as soon as the camera or subject moves; set type and place logos in the edit.
- Make it sharp. No motion blur, no heavy compression, no upscaling artifacts; export PNG or a high-quality JPG.
- Leave room for the move. Frame a little wider than the final crop if the camera will push in, and leave space on the side the subject will walk or turn toward.
- Choose a pose that can start moving. Weight on one foot, a hand mid-gesture, a head about to turn. A peak moment has nowhere to go.
- Get hands right or keep them out. The model will move whatever hands are in the frame, extra fingers included.
- Give the light a clear direction the clip can continue, and leave lens flares for the edit.
- For products, turn the label to camera, keep it large enough to read, and include a contact shadow so the product sits on something.
- Check the frame against your look bible, so shot 7 cuts next to shot 6.
Prompt for the motion
The vendors give the same advice in different words. Google's Veo guidance says "Focus your prompt on the motion you want to see," warns against trying to "Re-describe the character, the background, or the lighting depicted in the image," and says to "Prompt for camera movement, subject animation, and environmental changes." Runway's guide says "Effective image to video prompts focus almost exclusively on motion" and offers a pattern, "The camera [motion description] as the subject [action]." BytePlus sums up Seedance prompts as "subject + motion, background + motion, camera + motion." Kling warns that "A description that significantly deviates from the image may cause a camera cut or transition." Google's guide also tells you to avoid quotation marks in the prompt.
A prompt for a five-second product shot might read like this example.
Camera: Slow push in at eye level, no cuts. Subject: She lifts the can from the counter and pauses before drinking. Product: The label stays facing camera and unchanged. The logo does not move or morph. Environment: The curtain behind her stirs slightly. The window light stays constant. Pace: Slow and even. The move ends at a medium close-up.
One camera move and one action per clip is a good rule. If a shot needs two moves, make two clips and cut between them.
When to add a last frame
A last frame tells the model where the shot must end. It helps most for reveals and transformations: a closed box that ends open, a product that ends in a hand, a push in that ends on the clean plate for the end card. Veo 3.1, Kling 3.0, Luma Ray3.2 and MiniMax's Hailuo accept one. Runway's Gen-4.5 does not, and handles keyframes in a separate Animate Frames app.
Make both frames from the same look bible and the same light, so the model animates a move and not a change of scene, and keep the distance between them small enough to cross in the clip length you chose. Kling will not generate from an end frame alone, and Luma does not allow start and end frames on 10-second clips.
Check the contest rules on stills
Tool-locked contests do not all treat stills the same way. The Formula E Creators Challenge, which closes September 30, 2026, requires 100% of the video to be generated with Google models, but its rules also say "Stills or audio may use any tool." Other contests require the whole piece to come from one platform, and Runway's Big Ad Contest rules required entries to "Be created using tools available within the Runway platform." Read the rule for your contest before you make first frames in a different tool, and say which tools made the stills in your submission notes. For festivals that take work from specific tools, see our lists for Seedance users and Veo users.
Make the frames in Overs, animate them anywhere
Overs makes exactly the kind of still this page asks for. It studies the brand, casts a model when the product is worn or held, plans the shots, and gives each photo only the reference pictures it needs. You choose the image model that makes the final photos. Nano Banana 2 is the default and holds products steady across a batch, GPT Image 2.5 Sunburst is the slowest and most exact at keeping a product identical while the scene changes, and GPT Image 2 is often better at text, fabric and bright studio light.
Two features line up with the first-frame rules above. Overs sets headlines, sublines and buttons in a separate typography step and keeps the clean version, so you always have a text-free frame to animate. And it writes a text motion prompt for each photo for Seedance or Veo, which you take with the frame to your video tool, since Overs does not make video.
The numbers make the case for fixing frames before motion. Across 361 real renders from four 2026 campaigns by 100 Creatives, the agency Abhi Chawla founded, it took about 8 tries to get a usable photo. At published prices of $0.034 to $0.211 a render, that is $0.27 to $1.69 in AI fees per usable frame, before you spend a cent on video. Overs' guide to making AI product photos look real covers the ten tells that also ruin a first frame.
Overs is a sister product to AI Film Contests, from the same founder, and Abhi's own businesses make their images with it. You see the shot plan and an estimated cost before anything renders, and the free plan covers 40 photos a month on your own OpenRouter key, with no markup from Overs. Try Overs free at www.overs.studio and build your next spot from frames you have already approved.
Frequently Asked Questions
What makes a good first frame for image-to-video?
A good first frame is sharp, at or above the output resolution, in the same aspect ratio as the video, free of text and logos, and composed with room for the motion you plan. Hands should be correct or out of frame, the light should have a clear direction, and a product's label should face camera. It should also match the film's look, so the shot cuts with its neighbors.
Should the prompt describe the image when I use image-to-video?
Describe the motion. The image already tells the model what is in the scene, so spend the prompt on the camera move, the subject's action, the pace and what must stay still, such as a product label. Google, Runway, ByteDance and Kling all say this in their own guides, and Kling warns that a prompt that strays far from the image can cause an unwanted cut.
What aspect ratio should the first frame be?
The same as the video you want. Runway Gen-4.5 accepts 1:2 to 2:1 and crops anything else from the center, Kling 3.0 accepts 1:2.5 to 2.5:1, Seedance accepts ratios between 0.4 and 2.5 and center-crops a mismatch, and Google says Veo inputs of other sizes may be resized or cropped, advising 16:9 or 9:16 at 720p or higher. Crop the still yourself so no tool decides the framing for you.
Can I make first frames in one tool and animate them in another?
Usually yes, but check your contest's rules first. The Formula E Creators Challenge requires all video to come from Google models while allowing stills from any tool. Other contests require every part of an entry to come from their own platform. Say which tools made the stills and which made the video in your submission.