Your Voice Agent Doesn’t Have a Model Problem. It Has a Wiring Problem.
When a voice agent feels laggy, the instinct is to swap in a faster model. In a typical stitched pipeline, the model is rarely where the time goes — the network hops between separately-hosted speech, reasoning and synthesis services are. Here is the 800-millisecond budget, spent honestly.
- Subject
- Voice AI Latency: The 800 ms Budget, Itemised
- Published
- 19 AUG 2026
- Reading time
- 9 min
- Film
- Watch · 4:11
In this post · 4 sections
A voice agent starts feeling laggy, and the first instinct is almost always the same: swap in a faster model. It is usually the wrong layer. In a typical stitched voice pipeline — speech recognition from one vendor, the language model from another, speech synthesis from a third — the model call is rarely the only large thing in the budget, and it is almost never the thing worth attacking first. The network hops between those three separately-hosted services cost about as much as the model call itself — and unlike the model, they are entirely a wiring decision. Latency, in most voice agents that feel slow, is a systems-integration problem wearing a model-selection costume.
Why 800 milliseconds is the number
It is the budget we design to, and it comes from what a conversation sounds like rather than from what a pipeline can manage. Below roughly 800 milliseconds of response time, a spoken exchange still feels like a conversation. Past roughly 1,500 milliseconds, callers reliably describe the interaction as broken — not slow, broken, the way a phone call with a bad satellite delay feels wrong in a way that’s hard to name but impossible to miss. That window is what we built our own voice agent architecture to: streaming speech recognition, an agent runtime that can call out to CRM, scheduling and payment systems mid-conversation, all inside an 800-millisecond round-trip budget.
LiveKit, “Understand and Improve Voice Agent Latency” — on the multiple compounding sources of pipeline latency and why per-stage span breakdowns, not an aggregate number, are what actually locate the faulty stage. livekit.com/blog/understand-and-improve-agent-latency
The four-layer budget, and why it’s additive
A naive pipeline spends time in four places, sequentially: speech-to-text, the language model’s first token, text-to-speech’s first audio chunk, and network transport between all of them. Converging estimates from independent voice-AI engineering write-ups put typical component ranges at roughly 100–300ms for STT, 350–1,000ms for the LLM, 90–200ms for TTS, and 50–200ms of network transport — presented here as converging industry estimates, not one authoritative benchmark. The chart below picks one representative point inside each of those ranges to make the shape concrete; it is a composed illustration, not a trace from a specific call.
The four stages are additive in a naive architecture because each hop between separately-hosted services is its own network round trip, stacked on top of that service’s own processing time. That is the entire difference between the two bars above: the same speech recognizer, the same language model, the same speech synthesizer, wired two different ways. Nothing about swapping any one service for a faster competitor would have closed that gap — the gap was never inside any one component, it was between them.
What changes when the model calls out mid-conversation
Our own architecture’s agent runtime doesn’t just talk — it looks up a customer record, checks a calendar, and can move money, in the middle of a live call. The moment a conversational turn includes a tool call to a CRM, scheduling system or payment processor, that call is part of the same latency budget, not a separate concern with its own more relaxed expectations. A caller does not experience “the agent is now doing an API call” as a different category of wait than “the agent is thinking” — it is all just silence, measured against the same 800 milliseconds.
Mitigations, ranked by actual leverage
- Colocate the pipeline stages. The single highest-leverage fix, and the one the chart above is entirely about: put STT, the model, TTS and your tool integrations in the same region — ideally the same network — so the hop between them is single-digit milliseconds instead of a full round trip to another provider’s infrastructure.
- Stream partial responses. Don’t wait for a complete model turn before starting to speak it. Streaming the first coherent phrase to TTS while the model is still generating the rest overlaps two of the four stages instead of running them in sequence.
- Make TTS interruptible. A caller who starts talking over the agent should cut the agent off immediately, not after the current sentence finishes rendering. This doesn’t reduce the budget on paper, but it’s most of what “feels responsive” actually measures in practice.
- Only then, model choice. Once the wiring is colocated and streaming, a faster model genuinely helps close the remaining gap. It is the last lever, not the first one, because on a naive, cross-vendor pipeline it is fighting a problem it cannot solve.
None of this is a case against using best-in-class vendors for STT, the model, or TTS individually. It’s a case for treating the wiring between them as a first-class part of the architecture, instead of an afterthought that only gets attention once the obvious fix — a faster model — has already failed to help.
Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.