Video Generation

Four Image-to-Video Model Lines, One API: What Actually Differs

OpenRouter's comparison shows duration, frame roles, and per-second pricing vary more than model names — here's how to pick.

Four Image-to-Video Model Lines, One API: What Actually Differs — article cover
On this page6 SECTIONS
  1. One endpoint, very different contracts
  2. Price is per second, so duration is a budget line
  3. Two fields that decide what your image does
  4. The retry trap
  5. Takeaway
  6. Sources

If you already have the image your video should start from, the model name matters less than the settings behind it. OpenRouter’s comparison of image-to-video models, published September 29, 2026, makes that concrete by putting four model lines — Veo 3.1, Seedance, Kling, and Grok Imagine Video — into one table with the same columns: resolution, duration, frame roles, audio, and price per second.

The interesting part isn’t the specific numbers. It’s that all of these models share one submission lifecycle and diverge almost entirely on constraints you can check before spending a cent.

One endpoint, very different contracts

Every model in the comparison goes through POST /api/v1/videos, with the same submit, poll, and download flow. What changes per model is the shape of the job you can ask for:

  • Veo 3.1 variants generate 4, 6, or 8 second clips with first- and last-frame control and optional audio; Veo 3.1 and Fast reach 4K, Lite stops at 1080p.
  • Seedance 2.5 runs 4 to 30 seconds at 480p or 720p and accepts up to 50 reference assets — the only line here producing a 20–30 second clip in one generation.
  • Kling v3.0 covers 3 to 15 seconds at 720p with first/last frame control; audio toggles the price.
  • Grok Imagine Video 1.5 supports a first frame only, so you can pin the opening image but never the closing one.

Before comparing anything else, OpenRouter’s advice is to narrow the list by duration and frame control. That’s the right order of operations: if your clip must end on a second exact image, Grok Imagine is out regardless of price. If you need 25 seconds, Veo’s 8-second ceiling is out.

Price is per second, so duration is a budget line

Most models bill per second, which turns your storyboard into a cost estimate. At Seedance 2.5’s starting rate of $0.1028/s, a 4-second clip starts around $0.41 while a 30-second clip starts around $3.08. Veo 3.1 Lite starts at $0.03/s — roughly 6x cheaper than full Veo 3.1 at $0.20/s — which is why OpenRouter’s own cookbook uses Lite while prompts and motion are still being iterated.

Two pricing gotchas worth flagging:

  • Seedance bills per video token, not per second. The model pages derive seconds as height × width × duration × 24 ÷ 1,024; the per-second figures are starting rates for 480p. Higher resolutions change the math.
  • Kling charges differently with audio on: Standard jumps from $0.084 to $0.126 per second. Audio is a pricing decision, not just a feature flag.

I covered the same principle in Picking an Image-to-Video Model: Price Per Second Beats Model Hype — estimate from rate × duration before you touch a leaderboard.

Two fields that decide what your image does

The API separates two image roles. frame_images pins an image to the exact first frame, last frame, or both. input_references provides material — images, video, or audio — the model should follow without locking it to a frame position. If you send both, frame_images wins and the job counts as image-to-video.

This distinction is easy to get wrong. A character-consistency shot probably wants references; a product video that must open on the actual product shot needs frames. Not every model supports both roles, and GET /api/v1/videos/models exposes the supported_frame_images array so you can verify rather than guess.

One caveat OpenRouter is careful about: “Not listed” for reference inputs means the model description doesn’t mention them — not that references were tested and rejected.

The retry trap

The most operational detail in the post: the TypeScript SDK retries 5XX responses with exponential backoff for up to an hour by default. For video submission, that’s dangerous. A submission starts billable work; if the server accepts the job and a transient error hits before you receive the job ID, an automatic retry submits a second paid job. OpenRouter passes { retries: { strategy: "none" } } on the submission call so failures surface as errors you can act on deliberately.

That’s a pattern worth copying wherever an API call triggers billing on the server side.

Takeaway

Pick by constraints, not brand: fix duration and frame requirements first, then compare audio and per-second cost on whatever survives. And check the model page before estimating production cost — prices here were verified September 11, 2026, and vary by resolution, audio, and input type.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

Found this useful?

Support more practical AI articles, tutorials, and build notes.

Buy us a coffee
SHAREXEMAIL