Blog · For developers
Latency in voice AI: why a second feels like forever
On a phone call there's no spinner. Silence is the only loading state, and people read it as a fault.
By the CosVoice team · · 5 minute read
Count it out loud
Try this with someone you know. Ask them a question, and make them count "one Mississippi" in their head before they answer. It's unbearable. They look confused, you feel rude, and the conversation limps for the next minute. People start replying almost as the other person finishes, sometimes before. We don't notice that we do it. We notice at once when someone doesn't.
A phone call is worse than face to face, because there's nothing to look at. No typing indicator, no face thinking. On a screen a short wait is a wait. On the phone it's "hello? are you there?"
Four places the time goes
When a caller stops talking and your agent starts, the gap between is the sum of several smaller gaps. Measure only the model and you'll miss most of them. A speech-to-speech model folds the middle two together, which removes some hand-offs, though the work still has to happen somewhere.
- The network. Audio travels from the caller's phone through the carrier to your media server, and the reply makes the same trip back. Audio moves in small chunks, and every hop holds on to it for a moment.
- Speech in. Something has to decide the caller has finished. In a chained system, speech-to-text then has to settle on its transcript.
- Thinking. The language model reads the conversation so far and starts a reply. What counts is how soon the first words come, not how long the whole answer takes.
- Speech out. Text-to-speech turns those first words into audio, and the audio goes back down the line.
The pause you add on purpose
The sneaky one is inside speech in. How do you know the caller has finished? Mostly by waiting to see if they say anything else. That wait is a setting, which makes it a delay you chose.
Set it short and the agent jumps in whenever someone stops for breath. People read phone numbers aloud in groups with gaps between, and an eager agent will barge in between the area code and the rest. Set it long and every reply starts late, even the easy ones.
Better end-of-turn detection looks at the words as well as the silence. "My number is" is plainly not the end of a thought, however long the gap after it. If you're building this yourself, it's one of the places where effort pays back most.
What callers notice
Callers don't perceive milliseconds. They perceive behavior. We'd rank what they react to roughly like this, worst first.
- Being cut off. Worse than slowness. An agent that talks over someone has told them it isn't listening.
- A late first reply. The gap after their hello sets the tone for the whole call. A slow first word sounds like a dialer.
- Uneven gaps. A steady rhythm, even a slightly slow one, is easier to talk to than quick replies with the odd long stall.
- A voice that won't stop. They try to cut in and it carries on to the end of the paragraph.
Help that doesn't need a new model
Stream everything. Start speaking as soon as the first sentence exists, and don't wait for the whole reply. Keep that first sentence short so it's ready sooner. Put your media servers near your carrier's. Keep connections open between turns so you're not setting them up again mid-call.
Move slow work off the conversational path. If the agent has to look something up mid-call, it can say so, the way a person says "let me check". That's honest, and callers accept it.
What we'd avoid is canned filler on every turn. An "mm, good question" before each answer buys you a moment and costs you the caller's trust by about the third one.
Quicker, or interruptible
Here's the trade-off. A voice that can be interrupted has to keep listening while it speaks, and be ready to stop and drop what it was about to say. That's more work per turn than a voice that talks, finishes, then listens. The more natural-sounding voices tend to be the heavier ones as well.
We make that choice visible on our own line. There are 38 voices in two sets. The new voices can be interrupted and sound more natural, and they take a little longer to answer. The classic voices are the clearest on a phone line and the quickest to reply, but they finish the sentence before they listen. You pick in Settings, or your agent changes it with update_settings, and it applies to the next call.
Which one? For back-and-forth where people cut in, a new voice. For a busy front desk where clarity and speed matter most, a classic one. The longer argument is in full duplex vs turn-based voice.
Measure your own calls
We're not going to give you numbers, because the ones that matter are yours. A figure measured on a good connection with a short prompt tells you little about a caller in a van on a mobile.
So log a time at each stage of a turn: when the caller stopped, when you decided they'd stopped, when the first words of the reply existed, when audio left your server. If you use a finished line like ours, you can't see inside the turn, but each entry in the transcript from get_call carries a time, which is enough to see the rhythm of a call. How an AI phone call works covers the rest of the path.
Then look at the slowest turns, not the average. A call with nine quick replies and one long stall is remembered for the stall.
Common questions
What causes latency in voice AI?
Four things added together: the network trip, working out that the caller has finished and turning their speech into something the model can read, the model starting its reply, and turning that reply into audio. The wait to confirm the caller has stopped is the one people overlook.
How fast does a voice agent need to respond?
Fast enough that the caller doesn't say hello twice. There's no single number that fits every call, and a steady rhythm matters as much as raw speed. Test on a real phone on a mobile network, not on a laptop.
Will a faster language model fix voice AI latency?
Only the thinking part. If the delay is in end-of-turn detection, the network or speech output, a faster model won't change what the caller hears.
Why do interruptible voices take longer to answer?
They listen while they speak and have to be ready to stop mid-sentence, which is more work per turn. On our line the voices that can be interrupted take a little longer to answer than the classic ones.
Read next: Can your caller interrupt? · Choosing a voice for your line · Hear full calls
Hear it for yourself.
Call our own AI agent and ask her anything, or get a number in your area code in about a minute.