Cloudflare

Set the Model to cloudflare/auto and Let the Gateway Do the Picking

Cloudflare's AI Gateway Auto Router picks a capable-enough model per request, cutting internal coding spend by up to 30%.

Set the Model to cloudflare/auto and Let the Gateway Do the Picking — article cover

Budgets and dashboards tell you who is overspending. They don’t stop it. Cloudflare’s new Auto Router, released September 30, 2026 in public beta through AI Gateway, takes the next step: the gateway itself picks the model. Set the model name to cloudflare/auto, and each request gets routed to something capable enough for the task instead of defaulting to whatever frontier model the user typed.

Cloudflare reports up to 30% savings in its own OpenCode deployment compared to running everything on frontier models like OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5.

Why the savings exist at all

Most teams already understand that summarizing an email thread doesn’t need Opus. The problem is that individual users pick models manually in harnesses like Claude Code, Codex, and OpenCode, and nobody wants to be the person blocking a capable model from the security team. So everyone defaults high.

Cloudflare frames the opportunity as a “jagged frontier”: for any given task, some model in your portfolio can solve it, and the router’s job is finding which one while balancing quality and price. Savings come from not paying frontier rates for non-frontier work.

The internal benchmark is worth reading closely. On 97 general knowledge-work tasks (email, calendars, files, finance) with three samples per task, cloudflare/auto scored 86.6% success at $2.10 total, versus GPT-6 Sol at 84.2% for $2.64 and Opus 5.5 at 96.6% for $5.91. Auto landed at 80% of Sol’s cost and 35% of Opus’s, with a success rate between them. That’s the honest tradeoff: you’re buying most of the frontier quality at a fraction of the price, not all of the quality for free.

The routing logic, briefly

When a request hits cloudflare/auto, the gateway first filters the model pool by what can actually serve it: request format, credentials, spend limits, and upstream health. Then a multi-head classifier running on Workers AI scores the conversation across 14 task categories and four dimensions — complexity, ambiguity, stakes, and context dependence. A scoring matrix combines those signals with benchmark-derived weights and token prices to pick the highest-utility model.

Two details stand out for anyone who has run long agent sessions:

  • Cheaper list price doesn’t mean cheaper outcome. A model can burn disproportionately more tokens per task, so the router optimizes predicted trajectory cost, not dollars per million tokens.
  • Switching models has a cache cost. Mid-session switches throw away a hot prompt cache and force a full context rewrite, and most models can’t read another model’s reasoning tokens. The router applies a switching penalty that grows with context depth, so the deeper the session, the more a switch has to earn back.

The architecture is also legible — you can inspect a task’s predicted category and difficulty and trace why a model was chosen — and adding a new model means adding weights, not retraining.

What this changes for your stack

The broader shift is that the gateway moves from observing and enforcing to deciding. Budgets, spend analytics, and identity-aware limits have been the standard tooling; they still rely on individuals making cost-conscious choices request by request. Auto Router pushes that decision into infrastructure, which is where it belongs. This fits the argument from our earlier post on Sonnet 5.5’s cost tiering — the meaningful savings live in matching model capability to task, not in chasing headline prices.

Practical caveats before you adopt it. The benchmark covers general knowledge work; Cloudflare says results on coding are comparable to frontier models but your workload will differ. Cloudflare plans a cloudflare/auto-best profile that drops the cost tradeoff entirely, plus planned work on reasoning-level selection, provider capacity awareness, and zero-data-retention filtering. It’s free during beta, so the cheapest way to evaluate it is to point a non-critical internal workload at cloudflare/auto and compare your own cost-per-success numbers against your current default.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

Found this useful?

Support more practical AI articles, tutorials, and build notes.

Buy us a coffee
SHAREXEMAIL