The Cheapest Model That Clears Your Eval Bar Is the Correct Model
Most teams pick a frontier model in the kickoff meeting and never revisit it. That is not a technology decision — it is the absence of one. Here is what a routing harness actually looks like, and the three places a cascade quietly breaks.
- Subject
- LLM Model Routing: The Cheapest Model That Passes
- Published
- 19 AUG 2026
- Reading time
- 9 min
- Film
- Watch · 1:49
In this post · 4 sections
Ask most engineering teams why they use the model they use, and the honest answer is a kickoff meeting from six months ago. Someone picked a frontier model because it was the best one available at the time, wired it into the codebase, and it has not come up again since. That is not a technology decision. It is the absence of one — and it is quietly the most expensive line item most AI products carry, because a single fixed model is being asked to handle everything from “summarize this in one sentence” to the one query in a thousand that actually needs its full reasoning budget.
The fix is not a better model. It is a harness that decides, per query, which model is worth paying for — and most teams simply do not have one.
What a cascade actually buys you
The idea is not new. In 2023, a Stanford team published FrugalGPT, showing that routing queries through a cascade of models — try the cheapest first, escalate only when its answer looks unreliable — could match a frontier model’s accuracy at up to a 98% cost reduction, or beat it by a few points at the same total cost.
Chen, Zaharia & Zou, “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance,” arXiv:2305.05176 (2023). arxiv.org/abs/2305.05176
What has changed since 2023 is not the mechanism — it is the incentive. The price spread between a fast model and a frontier one has widened, not narrowed. In August 2026, Anthropic’s published API pricing put the fast rung of its Haiku/Sonnet/Opus ladder at roughly a fifth the input and output cost of that ladder’s top rung — and the same lineup now carries tiers priced above Opus, so the spread between cheapest and dearest is wider still, which only sharpens the argument for routing.
Anthropic API pricing, August 2026. claude.com/pricing
A five-times spread means a cascade that resolves most of its traffic at the cheap tier is not a rounding error on your inference bill. It is the difference between a unit economics problem and a solved one. That gap is the entire argument for building a harness instead of picking a model.
The anatomy of a routing harness
A cascade is not “try cheap, catch errors.” Done properly it has four distinct parts, and skipping any one of them is where teams that attempt this quietly fail:
- A golden set. A labelled collection of real queries with known-good answers, covering the actual distribution of what your system sees in production — not a handful of happy-path examples written the day before launch. This is the thing that turns “does the cheap model seem fine” into a number you can gate a release on.
- A confidence signal. The model’s own stated confidence is close to useless — language models are not well calibrated about their own uncertainty. Better signals are external: agreement between two independent samples, whether the output validates against an expected structure or schema, or a small, cheap grading pass that checks the answer against the query rather than trusting the model that produced it.
- An escalation rule. A concrete threshold on that signal, plus a decision about what happens at the boundary — retry at the next tier, flag for human review, or refuse. “Escalate when unsure” is not a rule until someone has written down what “unsure” means as a number.
- The right metric. Cost per resolved query — not cost per call. A cascade that fails at the cheap tier 40% of the time is paying for that tier’s calls and the escalation on top of them. If your accounting only tracks call volume by tier, a badly-tuned cascade can look cheaper than it is.
Where cascades break
Three failure modes show up often enough to name:
- Latency-sensitive paths. Every escalation is a sequential round trip. That is a non-issue for a batch job and a real problem for anything with a tight interactive budget — a voice agent with an 800-millisecond turn-taking window does not have time to fail cheap and retry expensive. Cascades belong where you can afford the tail latency of the worst case, not everywhere.
- Tasks with no clean definition of “good enough.” A confidence threshold needs something to be confident about. Open-ended creative or judgment-heavy tasks often have no ground truth to calibrate against, which means the threshold is measuring the model’s fluency, not its correctness — and fluency is exactly what a cheap model is often best at faking.
- Confidently wrong at the cheap tier. Adversarial or out-of-distribution inputs can fool a confidence signal into reading high when the answer is wrong. This is the failure mode that should worry you most, because it does not look like a failure — it ships fast and wrong instead of slow and right, and nothing downstream flags it.
What actually holds up in production
A harness worth running is never “set once, forget.” Two things keep it honest. First, the golden set grows from production, not just from launch-day test cases — every time a human overrides the cheap tier’s answer, that case becomes a new eval example, so the bar gets harder to clear over time rather than staying frozen at day-one difficulty. Second, the whole cascade gets re-evaluated on a schedule, not left alone once it works. A provider refreshing a model behind the same name can silently move the quality/cost frontier your thresholds were tuned against — the harness is the only thing that would tell you that happened before a customer does.
None of this is worth building for a handful of calls a day. It earns its cost once you have enough query volume that the price spread between tiers is real money, and enough repetition in the query shape that a golden set can actually represent what you see. Below that line, pick whichever model clears the bar by inspection and move on — a routing harness you don’t need is just another system to maintain. Above it, the model you use should be an output of the harness, not an input to the architecture.
Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.