If you route requests through an auto-router, you’re trusting an opaque middle layer with your quality, latency, and bill. On October 2, 2026, OpenRouter launched Model Router Benchmarks to make that layer comparable: every listed router was run through benchmarks from varied domains and scored on quality, speed, and cost, with individual models shown alongside as a baseline.
Why routing is tempting — and risky
The economics are real. OpenRouter’s announcement compares DeepSeek v4 Flash with GPT-6 Astra: the latter’s average price per token is over 48x higher, and running an average 10-to-49-turn session in Codex costs 21x more. Routers try to capture those savings by shifting parts of a task to cheaper models, switching based on task type (legal research vs. coding), or adjusting model intelligence to the stage of an agentic session.
But the announcement is refreshingly honest about the failure modes. Switching models can rebuild your input cache, making requests more expensive. A router often can’t judge task complexity from the prompt alone. Heuristics about session stage may disagree with what the model is actually about to do. And every extra routing decision adds latency. Sometimes a single well-chosen model just wins.
The benchmarks exist, per OpenRouter, to show which routers actually overcome these challenges — so you can decide when, and whether, to adopt one.
Not all routers mean the same thing
The page also untangles a term that gets thrown around loosely. Provider routing — choosing which inference provider serves a given model — is what OpenRouter has done since day one. Model routing is the newer thing: you send a request to a router and it picks the model, behaving like a model itself.
Among the benchmarked model routers, three approaches stand out:
- Blend executors. Unbiased’s Pareto and Sakana’s Fugu blend models internally, sometimes escalating to frontier models, without disclosing which models run or how often they switch. Billing is standardized per token.
- Per-turn selectors. Auto Router and Jev Router pick a model each turn, and the selection is visible — you pay standard rates for whichever model was chosen.
- Pair swappers. NVIDIA’s Switchyard pairs a cheaper model with a stronger one and alternates as an agent works. OpenRouter benchmarked several pairs, but you can run any pairing you like.
Some styles were excluded because they serve specialized purposes: Fusion runs one request across a council of models and synthesizes an answer, while single-model aliases like Pareto Code and the -latest slugs don’t involve real selection. That’s a fair boundary — comparing them on the same axes would mislead more than inform.
How to read the score
Comparing across quality, speed, and cost is apples-to-oranges, so OpenRouter’s Router Index folds them into a 0-to-10 score per benchmark. Default weighting is 60% quality, 20% time per task, 20% cost — and there’s a slider to rebalance it if your priorities differ. For a cost-sensitive support bot, that slider alone changes the leaderboard significantly.
Adoption is intentionally low-friction: all listed routers are available on OpenRouter, and you swap one in wherever you’d put a model string — app, harness, or API call.
What I’d actually do with this
The most useful framing here is the baseline. Because individual models sit alongside the routers, the page answers the question that matters: does this router beat just picking a good model myself? If a router’s index score barely clears a well-chosen model, the opacity and switching costs probably aren’t worth it. I’ve made a similar distinction before between routing and orchestration in Three Layers, One Buzzword: Where Routing Ends and Orchestration Starts — this page is essentially OpenRouter giving that middle layer a report card.
Keep two caveats in mind. OpenRouter says the benchmarks reflect general tasks, not your workload, so treat the scores as a shortlist generator rather than a verdict. And the scores will keep changing — OpenRouter commits to adding routers and updating numbers over time. The practical next step: pick your two or three heaviest request patterns, re-run them through the top-scoring router on your preferred index weighting, and compare against your current single-model baseline before committing.
