The Cheapest Model That Clears Your Eval Bar Is the Correct Model

Most teams pick a frontier model in the kickoff meeting and never revisit it. That is not a technology decision — it is the absence of one. Here is what a routing harness actually looks like, and the three places a cascade quietly breaks.

Subject
LLM Model Routing: The Cheapest Model That Passes
Published
19 AUG 2026
Reading time
9 min
In this post · 4 sections
  1. 01What a cascade actually buys you
  2. 02The anatomy of a routing harness
  3. 03Where cascades break
  4. 04What actually holds up in production
The same argument in 1:49. Our graphics, AI narration.

Ask most engineering teams why they use the model they use, and the honest answer is a kickoff meeting from six months ago. Someone picked a frontier model because it was the best one available at the time, wired it into the codebase, and it has not come up again since. That is not a technology decision. It is the absence of one — and it is quietly the most expensive line item most AI products carry, because a single fixed model is being asked to handle everything from “summarize this in one sentence” to the one query in a thousand that actually needs its full reasoning budget.

The fix is not a better model. It is a harness that decides, per query, which model is worth paying for — and most teams simply do not have one.

What a cascade actually buys you

The idea is not new. In 2023, a Stanford team published FrugalGPT, showing that routing queries through a cascade of models — try the cheapest first, escalate only when its answer looks unreliable — could match a frontier model’s accuracy at up to a 98% cost reduction, or beat it by a few points at the same total cost.

Chen, Zaharia & Zou, “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance,” arXiv:2305.05176 (2023). arxiv.org/abs/2305.05176

What has changed since 2023 is not the mechanism — it is the incentive. The price spread between a fast model and a frontier one has widened, not narrowed. In August 2026, Anthropic’s published API pricing put the fast rung of its Haiku/Sonnet/Opus ladder at roughly a fifth the input and output cost of that ladder’s top rung — and the same lineup now carries tiers priced above Opus, so the spread between cheapest and dearest is wider still, which only sharpens the argument for routing.

Anthropic API pricing, August 2026. claude.com/pricing

A five-times spread means a cascade that resolves most of its traffic at the cheap tier is not a rounding error on your inference bill. It is the difference between a unit economics problem and a solved one. That gap is the entire argument for building a harness instead of picking a model.

A three-tier routing cascade: a query tries the cheapest model first and escalates only when a confidence gauge fails to clear the barROUTINGTIERED BY CONFIDENCEESCALATE ONLY ON SIGNALDWG Nº 01COST/RESOLVED QUERYQUERYRESOLVEDshipped tothe callerTIER 1Fast & cheap$CONFIDENCE BAR↓ escalateTIER 2Balanced$$BAR↓ escalateTIER 3Frontier$$$BARlast resort, always clearsSHAPE OF THE TRAFFIC — ILLUSTRATIVE, NOT A MEASURED RESULTTIER 1most queriesTIER 2escalatedTIER 3rareGolden set, not vibesEscalate on signal onlyCost per resolved query
How the cascade pays for itself: nearly everything clears at the cheap tier, a minority escalates once, and only the genuinely hard cases reach the frontier model — so the average cost per resolved query sits far below the frontier price, without ever shipping a wrong answer the eval bar would have caught.

The anatomy of a routing harness

A cascade is not “try cheap, catch errors.” Done properly it has four distinct parts, and skipping any one of them is where teams that attempt this quietly fail:

  • A golden set. A labelled collection of real queries with known-good answers, covering the actual distribution of what your system sees in production — not a handful of happy-path examples written the day before launch. This is the thing that turns “does the cheap model seem fine” into a number you can gate a release on.
  • A confidence signal. The model’s own stated confidence is close to useless — language models are not well calibrated about their own uncertainty. Better signals are external: agreement between two independent samples, whether the output validates against an expected structure or schema, or a small, cheap grading pass that checks the answer against the query rather than trusting the model that produced it.
  • An escalation rule. A concrete threshold on that signal, plus a decision about what happens at the boundary — retry at the next tier, flag for human review, or refuse. “Escalate when unsure” is not a rule until someone has written down what “unsure” means as a number.
  • The right metric. Cost per resolved query — not cost per call. A cascade that fails at the cheap tier 40% of the time is paying for that tier’s calls and the escalation on top of them. If your accounting only tracks call volume by tier, a badly-tuned cascade can look cheaper than it is.
Where this overlaps with security, not just cost: an escalation rule is also a place excess autonomy hides. A misconfigured threshold can silently route every query to the most expensive tier — or, worse, let a retry loop escalate itself with no ceiling. The OWASP GenAI Security Project’s guidance on excessive agency in LLM applications is written about tool permissions, but the same discipline — bound the blast radius, make the escalation path observable, put a hard ceiling on retries — applies directly to a routing layer. genai.owasp.org

Where cascades break

Three failure modes show up often enough to name:

  • Latency-sensitive paths. Every escalation is a sequential round trip. That is a non-issue for a batch job and a real problem for anything with a tight interactive budget — a voice agent with an 800-millisecond turn-taking window does not have time to fail cheap and retry expensive. Cascades belong where you can afford the tail latency of the worst case, not everywhere.
  • Tasks with no clean definition of “good enough.” A confidence threshold needs something to be confident about. Open-ended creative or judgment-heavy tasks often have no ground truth to calibrate against, which means the threshold is measuring the model’s fluency, not its correctness — and fluency is exactly what a cheap model is often best at faking.
  • Confidently wrong at the cheap tier. Adversarial or out-of-distribution inputs can fool a confidence signal into reading high when the answer is wrong. This is the failure mode that should worry you most, because it does not look like a failure — it ships fast and wrong instead of slow and right, and nothing downstream flags it.

What actually holds up in production

A harness worth running is never “set once, forget.” Two things keep it honest. First, the golden set grows from production, not just from launch-day test cases — every time a human overrides the cheap tier’s answer, that case becomes a new eval example, so the bar gets harder to clear over time rather than staying frozen at day-one difficulty. Second, the whole cascade gets re-evaluated on a schedule, not left alone once it works. A provider refreshing a model behind the same name can silently move the quality/cost frontier your thresholds were tuned against — the harness is the only thing that would tell you that happened before a customer does.

None of this is worth building for a handful of calls a day. It earns its cost once you have enough query volume that the price spread between tiers is real money, and enough repetition in the query shape that a golden set can actually represent what you see. Below that line, pick whichever model clears the bar by inspection and move on — a routing harness you don’t need is just another system to maintain. Above it, the model you use should be an output of the harness, not an input to the architecture.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.

Book an AI Architecture Review