OpenRouter

Cheap-First Routing for Support Bots: What Your App Owns vs. What the Gateway Handles

OpenRouter's new tutorial maps three app-side routing patterns for sending easy support tickets to cheap models and escalating the rest.

Cheap-First Routing for Support Bots: What Your App Owns vs. What the Gateway Handles — article cover

Most support bots overpay. A routine question like “how do I reset my password?” and a gnarly multi-system refund dispute do not need the same model, but plenty of teams route everything to the strongest one because it is simpler to wire. OpenRouter published a guide on October 2, 2026, on its blog’s tutorials page that tackles exactly this: cheap-first routing, where routine FAQ traffic goes to a small, inexpensive model and selected hard or uncertain requests escalate to a stronger one.

What the guide covers

Per the supplied RSS summary, the tutorial walks through three application-side routing patterns and — more usefully — draws a boundary between two layers of responsibility. Your application owns the routing logic: deciding which requests are routine, which are hard, and when to escalate. OpenRouter’s Auto Router and model fallbacks handle what happens downstream: picking among models and retrying when one fails.

That split matters when you are designing the system. If you assume the gateway will make escalation decisions for you, you will be disappointed — the gateway’s Auto Router and fallbacks cover availability and general capability, not your specific definition of “this question is too hard for the cheap model.” The tutorial includes a copyable request for a support flow, which is the practical part most routing discussions skip.

The metric that makes or breaks it

The guide also sets out how to measure two numbers: escalation rate and cost per resolved ticket. These are the right pair. Escalation rate tells you whether your cheap model is actually absorbing traffic — if nearly everything escalates, you have added latency and complexity without saving money. Cost per resolved ticket tells you whether the whole setup beats your baseline of routing everything to one model.

It is worth watching both together. A low escalation rate with a high unresolved-ticket rate is worse than routing everything to the expensive model, because you have traded correctness for cost and are measuring the wrong thing. The tutorial’s framing of “cost per resolved ticket” rather than cost per request quietly acknowledges this.

Three patterns, one decision

The summary does not detail the three routing patterns, so I will not invent them — but the existence of three distinct patterns is itself informative. It suggests the main design decision is not which model to use but where the routing decision lives in your application: before the request, inside the request, or on retry. Whichever you pick, the escalation rule is application logic you have to write and maintain.

This is a theme worth internalizing beyond support bots. I made a similar case in an earlier post on letting a gateway pick models with cloudflare/auto — delegation to infrastructure works well for general routing, but the decisions tied to your product’s quality bar stay with you.

A grounded takeaway

If you run a support bot, the cheapest improvement is often not a better model but a cheaper first attempt with a deliberate escalation path. OpenRouter’s guide gives you the patterns, a request to copy, and two metrics to track. The open question — and the part the summary does not resolve — is how to classify “hard or uncertain” requests reliably in practice, since that classifier becomes a piece of your product you now have to evaluate and maintain like any other component.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

Found this useful?

Support more practical AI articles, tutorials, and build notes.

Buy us a coffee
SHAREXEMAIL