Your Voice Agent Doesn’t Have a Model Problem. It Has a Wiring Problem.

When a voice agent feels laggy, the instinct is to swap in a faster model. In a typical stitched pipeline, the model is rarely where the time goes — the network hops between separately-hosted speech, reasoning and synthesis services are. Here is the 800-millisecond budget, spent honestly.

Subject
Voice AI Latency: The 800 ms Budget, Itemised
Published
19 AUG 2026
Reading time
9 min
In this post · 4 sections
  1. 01Why 800 milliseconds is the number
  2. 02The four-layer budget, and why it’s additive
  3. 03What changes when the model calls out mid-conversation
  4. 04Mitigations, ranked by actual leverage
The same argument in 4:11. Our graphics, AI narration.

A voice agent starts feeling laggy, and the first instinct is almost always the same: swap in a faster model. It is usually the wrong layer. In a typical stitched voice pipeline — speech recognition from one vendor, the language model from another, speech synthesis from a third — the model call is rarely the only large thing in the budget, and it is almost never the thing worth attacking first. The network hops between those three separately-hosted services cost about as much as the model call itself — and unlike the model, they are entirely a wiring decision. Latency, in most voice agents that feel slow, is a systems-integration problem wearing a model-selection costume.

Why 800 milliseconds is the number

It is the budget we design to, and it comes from what a conversation sounds like rather than from what a pipeline can manage. Below roughly 800 milliseconds of response time, a spoken exchange still feels like a conversation. Past roughly 1,500 milliseconds, callers reliably describe the interaction as broken — not slow, broken, the way a phone call with a bad satellite delay feels wrong in a way that’s hard to name but impossible to miss. That window is what we built our own voice agent architecture to: streaming speech recognition, an agent runtime that can call out to CRM, scheduling and payment systems mid-conversation, all inside an 800-millisecond round-trip budget.

LiveKit, “Understand and Improve Voice Agent Latency” — on the multiple compounding sources of pipeline latency and why per-stage span breakdowns, not an aggregate number, are what actually locate the faulty stage. livekit.com/blog/understand-and-improve-agent-latency

The four-layer budget, and why it’s additive

A naive pipeline spends time in four places, sequentially: speech-to-text, the language model’s first token, text-to-speech’s first audio chunk, and network transport between all of them. Converging estimates from independent voice-AI engineering write-ups put typical component ranges at roughly 100–300ms for STT, 350–1,000ms for the LLM, 90–200ms for TTS, and 50–200ms of network transport — presented here as converging industry estimates, not one authoritative benchmark. The chart below picks one representative point inside each of those ranges to make the shape concrete; it is a composed illustration, not a trace from a specific call.

A race between two versions of the same voice pipeline on one clock: the naive pipeline, with a 140-millisecond network hop between each separately hosted service, finishes at about 1,150 milliseconds, over the 800-millisecond budget; the integrated pipeline, with the same speech, model and synthesis stages and no hops, finishes at about 730 milliseconds, inside it.LATENCYSAME JOB, ONE CLOCKTHE HOPS BLOW THE BUDGETDWG Nº 03800MS ROUND TRIP800MS BUDGETONE CLOCK · BOTH LANES START AT 0NAIVE PIPELINE — SEPARATE VENDORSSTT120HOP140LLM450HOP140TTS100HOP140NET60≈1,150MS · DONE, OVER THE 800 BUDGETNET = LAST MILE · HOP = VENDOR TO VENDOR, 3 × 140MSINTEGRATED PIPELINE — COLOCATEDSTT120LLM450TTS100NET60≈730MS · DONE, INSIDE THE BUDGETFOUR STAGES INSIDE THE CITED RANGES · THE HOPS ARE OUR OWN ESTIMATEILLUSTRATIVE COMPOSITION, NOT A MEASURED TRACESame components throughoutOnly the wiring changedColocate first, model lastA race between two versions of the same voice pipeline on one clock: the naive pipeline, with a 140-millisecond network hop between each separately hosted service, finishes at about 1,150 milliseconds, over the 800-millisecond budget; the integrated pipeline, with the same speech, model and synthesis stages and no hops, finishes at about 730 milliseconds, inside it.LATENCYDWG Nº 03800MS ROUND TRIPSAME JOB, ONE CLOCKTHE HOPS BLOW THE BUDGETONE CLOCK · BOTH LANES START AT 0800MS BUDGETNAIVE — SEPARATE VENDORSSTT120HOP140LLM450HOP140TTS100HOP140≈1,150MS · OVER BUDGETINTEGRATED — COLOCATEDSTT120LLM450TTS100≈730MS · INSIDE BUDGETSTT 120 · LLM 450 · TTS 100 · NET 60 (LAST MILE)HOP = VENDOR TO VENDOR, 3 × 140MSFOUR STAGES INSIDE THE CITED RANGESTHE HOPS ARE OUR OWN ESTIMATEILLUSTRATIVE COMPOSITION, NOT A MEASURED TRACESame components throughoutOnly the wiring changedColocate first, model last
Same components, one clock: the naive pipeline's per-vendor network hops are the whole reason it misses an 800ms budget the integrated pipeline clears with room to spare — nothing about the model or the speech engines changed.

The four stages are additive in a naive architecture because each hop between separately-hosted services is its own network round trip, stacked on top of that service’s own processing time. That is the entire difference between the two bars above: the same speech recognizer, the same language model, the same speech synthesizer, wired two different ways. Nothing about swapping any one service for a faster competitor would have closed that gap — the gap was never inside any one component, it was between them.

What changes when the model calls out mid-conversation

Our own architecture’s agent runtime doesn’t just talk — it looks up a customer record, checks a calendar, and can move money, in the middle of a live call. The moment a conversational turn includes a tool call to a CRM, scheduling system or payment processor, that call is part of the same latency budget, not a separate concern with its own more relaxed expectations. A caller does not experience “the agent is now doing an API call” as a different category of wait than “the agent is thinking” — it is all just silence, measured against the same 800 milliseconds.

If you can’t see the four numbers separately, you can’t fix the one that’s actually broken. An aggregate “response took 1.4 seconds” log line tells you nothing actionable. Per-stage instrumentation — a span for STT, a span for first-token latency, a span for first-audio-chunk, a span for every outbound tool call — is what turns “it feels slow” into “the payment lookup added 340ms because it’s hosted in a different region.”

Mitigations, ranked by actual leverage

  1. Colocate the pipeline stages. The single highest-leverage fix, and the one the chart above is entirely about: put STT, the model, TTS and your tool integrations in the same region — ideally the same network — so the hop between them is single-digit milliseconds instead of a full round trip to another provider’s infrastructure.
  2. Stream partial responses. Don’t wait for a complete model turn before starting to speak it. Streaming the first coherent phrase to TTS while the model is still generating the rest overlaps two of the four stages instead of running them in sequence.
  3. Make TTS interruptible. A caller who starts talking over the agent should cut the agent off immediately, not after the current sentence finishes rendering. This doesn’t reduce the budget on paper, but it’s most of what “feels responsive” actually measures in practice.
  4. Only then, model choice. Once the wiring is colocated and streaming, a faster model genuinely helps close the remaining gap. It is the last lever, not the first one, because on a naive, cross-vendor pipeline it is fighting a problem it cannot solve.

None of this is a case against using best-in-class vendors for STT, the model, or TTS individually. It’s a case for treating the wiring between them as a first-class part of the architecture, instead of an afterthought that only gets attention once the obvious fix — a faster model — has already failed to help.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.

Book an AI Architecture Review