Most image-to-video model comparisons lead with quality screenshots. OpenRouter’s comparison, published September 29, 2026, takes a different angle: it treats the model choice as a constraints problem. If your clip has to open on an exact product shot, end on a second fixed image, run thirty seconds, or stay under a price per second, only some models qualify — and the differences are easy to check before you spend a credit.
Frame roles first, brand second
OpenRouter splits image inputs into two request fields. frame_images pins your image as the exact first frame, last frame, or both. input_references hands the model material to follow — a character’s appearance, say — without fixing it to any frame position. If you send both, frame_images wins.
This distinction matters for product teams. If the video must end on a specific closing card, only models supporting the last_frame role work. Per OpenRouter’s catalog, Grok Imagine Video 1.5 accepts first frames only, while Veo 3.1, Seedance, and Kling lines all support both roles. Each model’s supported_frame_images array in the catalog tells you this upfront, so it’s a lookup, not an experiment.
What the numbers actually say
The useful summary from the comparison:
- Veo 3.1 family: 4–8 second clips, up to 4K, first and last frame control, optional audio. Lite starts at $0.03/s; the full model starts at $0.20/s. Fast and Lite descriptions mention SynthID watermarking — relevant if you have provenance requirements.
- Seedance 2.5: the only model in the comparison producing 20–30 second clips in one generation, at 480p/720p, starting at $0.1028/s. It also accepts up to 50 reference assets. The 2.0 line caps at 15 seconds but starts cheaper.
- Kling v3.0: 3–15 seconds at 720p with both frame roles. Audio is optional and priced in — Standard is $0.084/s without it, $0.126/s with.
- Grok Imagine Video 1.5: 1–15 seconds, up to 1080p, first frame only, with generated audio baked in.
Two pricing gotchas are worth flagging. Seedance bills per video token (height × width × duration × 24 ÷ 1,024), not per second, so resolution changes the math directly. And since most models bill per second, clip duration is a cost decision: at Seedance 2.5’s starting rate, a 4-second clip runs about $0.41 while a 30-second clip starts near $3.08.
One API detail that can double your bill
The subtlest point in OpenRouter’s post is about retries. Their TypeScript SDK retries failed requests with exponential backoff by default — for up to an hour. But a video submission starts billable work the moment the server accepts it. If a transient error hits after acceptance but before you get the job ID, an automatic retry submits a second paid job and you only ever see one of them.
For submissions, they recommend passing retries: { strategy: "none" } and handling resubmission yourself. That’s a pattern worth copying wherever a request kicks off paid work: distinguish idempotent reads from billable writes at the retry-policy level, not the call site.
The same care applies to input handling. OpenRouter notes a URL that works in your browser can still fail when the provider fetches it — session cookies, HTML redirect pages, bot checks. A quick curl -I checking for a 200 and an image content type is cheap insurance before a paid generation.
This is the same discipline I wrote about in Descript’s eval loop — when every model trial has a real cost, structured selection beats ad-hoc experimentation.
The practical takeaway
Filter in this order: required clip duration, frame roles (does the clip need a fixed closing frame?), then audio, resolution, and price per second. For short audio-enabled clips, the Veo Lite tier is the cheap starting point; for long single-shot generations, Seedance 2.5 is currently the only option listed; for end-frame control at mid length, Kling v3.0 fits.
One caveat: OpenRouter checked these prices on September 11, 2026, and video pricing varies by resolution, audio, and input type — treat the figures as a comparison and verify on the model page before estimating production cost. The same lookup endpoint, GET /api/v1/videos/models, returns live supported settings, which is the smarter source of truth anyway.
Sources
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
