OpenRouter published a comparison of four image-to-video model lines on September 29, 2026: Google’s Veo 3.1 family, ByteDance’s Seedance line, Kuaishou’s Kling v3.0, and xAI’s Grok Imagine Video. The useful part isn’t the leaderboard framing — it’s that the requirements are narrow and checkable. If you know three things (clip length, whether the video must end on a fixed frame, and your price ceiling per second), the model list shrinks fast.
The three constraints that actually narrow the field
OpenRouter groups image usage into two request modes. frame_images pins an image to the exact first frame, last frame, or both. input_references gives the model material to follow without fixing it to any frame position. So before comparing models, ask: does this image have to appear at a specific point in the clip, or just guide the result? A product video that must open on an exact shot is image-to-video. A character whose look must persist across a new scene is reference-to-video.
Then duration and audio. Per the comparison, checked against the models catalog on September 11, 2026:
- Veo 3.1 / Fast / Lite: 4, 6, or 8 second clips, optional generated audio, first- and last-frame control. Veo 3.1 and Fast reach 4K; Lite stops at 1080p. Lite starts at $0.03/s.
- Seedance 2.5: 4 to 30 second clips at 480p or 720p, and it takes up to 50 image, video, and audio reference assets. It’s the only model in the comparison that produces a 20–30 second shot in one generation. Starting rate about $0.1028/s.
- Kling v3.0 Standard and Pro: 3 to 15 seconds at 720p, first and last frame control, audio optional. Standard is $0.084/s without audio, $0.126 with; Pro is $0.112 and $0.168.
- Grok Imagine Video 1.5: 1 to 15 seconds at up to 1080p, first frame only, with audio described as always generated. $0.08 to $0.25/s plus $0.01 per input image.
The practical split: if the clip must end on a second exact image, Grok Imagine Video is out — its supported frame roles list first_frame only. If you need 20+ seconds, Seedance 2.5 is the only one that fits. Everything else is a price and resolution trade.
Billing quirks worth reading before you estimate cost
Most of these models bill per second, which makes duration itself a cost lever — OpenRouter’s own math: at Seedance 2.5’s starting rate, a 4 second clip starts around $0.41, a 30 second clip around $3.08, before resolution and input changes. Two other wrinkles from the source:
- Seedance pricing is actually per video token (height × width × duration × 24 ÷ 1,024); the per-second figures are the starting rates shown at 480p. Requests that include video input use a lower per-token rate.
- Kling charges separately for audio, and Grok models add a per-input-image fee. If your pipeline auto-generates audio by default, you’re paying a hidden multiplier.
I’ve written before about why price per second beats model hype when picking a video model (that post is here), and this comparison is a good concrete instance: the “best” model at 4 seconds is often the wrong one at 30.
One API lifecycle, one submission gotcha
Every model here runs through the same POST /api/v1/videos submit-poll-download lifecycle, and GET /api/v1/videos/models returns each model’s supported settings and pricing SKUs. That’s the real benefit of a routing layer: swap model IDs without rewriting the job loop.
But there’s a submission trap OpenRouter flags that applies anywhere: their SDK retries 5XX errors with exponential backoff for up to an hour by default. Since a video submission starts billable work, a transient error after the server accepted the job can trigger a second paid job the client never notices. Their recommendation is to disable retries for the submission call specifically and resubmit yourself on failure. Also verify the image URL server-side (curl -I, expecting a 200 and an image content type) — a URL that works in your browser can still fail behind a session cookie or bot check.
A limitation to plan around
Seedance models are hosted-API only — OpenRouter found no published weights or license for any of them, so nothing runs locally or air-gapped. And a few reference-input fields are listed as “not described” on model pages, which per OpenRouter means unconfirmed, not rejected; worth testing before you build a pipeline on them. The cheapest first step: query the models endpoint, filter by your required duration and frame roles, then sort what’s left by price per second.
Sources
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
