AI Agent Technology · 3 September 2026 · 8 min read

Speech rhythm: why a fast AI voice agent can still sound slow

A millisecond figure measures the gap between turns. Callers hang up over the gap between their question and the answer, and that one is measured in whole seconds and whole minutes. Here is what to time on a demo call.

TIME TO A SUBSTANTIVE ANSWER90sLive front desk2sAgent, same accountagainstFive recorded calls on one client account: the front desk averagedabout 90 seconds to a substantive answer, the agent about 2 seconds.

Two clocks run on every phone call

Ask three voice-agent vendors how fast their agent is and you will get three figures in milliseconds. They are all measuring the same interval: the pause between the moment the caller stops talking and the moment the agent starts. That interval is real and it matters, which is why we wrote about the reply-speed threshold and about agents that talk over people in the first place. But it is a number about turns, and a caller does not experience turns. A caller experiences a call.

There is a second clock, and almost nobody times it. It starts when the caller asks the thing they actually phoned about and stops when they have the answer: the price, the slot, the yes or no. A vendor's slide reports the first clock. Customers hang up over the second one.

The two come apart more often than you would expect. An agent can start speaking the instant you finish and still take minutes to tell you a price, because it spends those minutes confirming what you said, asking something you already answered, and reading out a paragraph you did not request. Every turn is fast. The call is slow. That is a rhythm problem, and no further engineering on the first clock touches it.

What the second clock looks like on a real phone line

There is published evidence for how long people will let the second clock run. ContactBabel measured the average wait before a caller hears a human voice in UK contact centres, and how many callers give up before that happens. An older Velaro survey put a number on patience directly.

116 sec
average wait before a human voice, UK contact centres (ContactBabel, 2024)
8.4%
inbound callers who hang up before anyone answers (ContactBabel, 2024)
60%
will not hold past one minute (Velaro, 2012)

Those are figures about queues rather than about conversation, but they set the scale. The clock a caller is running is not measured in milliseconds. It is measured in whole seconds and whole minutes, and it runs out.

We have our own numbers on the same clock. On an anonymised client account at a hair-transplant clinic, our operators recorded five calls in which the same price objection was put to the client's live front desk and to a neuro-agent, and scored both. The front desk averaged about 90 seconds to a substantive answer; the agent averaged about 2 seconds. The longest hold before the front desk handed the question to a doctor ran 66 seconds. The agent needed 30 seconds to describe the scope zone by zone and 40 seconds to justify the price against a competing quote. Those are our own measurements on one client account, scored by us, and we do not offer them as research.

Look at those last two figures again. Thirty seconds and forty seconds are not latency. They are the length of a well-shaped answer: long enough to be worth having, short enough that the caller is still listening at the end of it. An agent that took several times as long to deliver the same content would have scored exactly the same on the first clock, and lost the call.

Four rhythm faults that survive a good latency figure

All four of these are compatible with an excellent millisecond number. All four are audible on a demo call once you know what to listen for.

01
The buried answer

The agent replies instantly and then talks past the point where the answer should have arrived. Everything it says is true; none of it is what was asked. Listen for where the answer sits inside the turn. A good one leads with it.

02
The confirmation loop

You give a date, the agent reads the date back, you confirm, the agent reads it back again with the time attached. Every turn is quick and the sequence is slow. Count the turns between giving information and the agent using it.

03
The re-asked question

Something you said early comes back as a question later. To a caller that is not slowness, it is not being listened to, and it costs more than a late reply ever does.

04
The unnarrated pause

The agent goes to look something up and says nothing while it does. Silence on a phone line reads as a dropped call, not as thinking. A receptionist fills it out loud, and so should an agent.

The rhythm we build to, and where the numbers come from

We do not write a call in milliseconds. We write it in blocks. The frame is seven of them: the open, earning the right to the conversation, reaching the right person, qualification, the offer, the objection, and a dated next step. The opening block runs 10 to 15 seconds, hello to the reason for the call. The first three blocks together run about 30 seconds. Qualification is two or three questions that change what gets offered, and takes a minute to a minute and a half. A cold call that is working reaches a dated next step in under four minutes.

Those figures are prescriptions rather than measurements, drawn from reviewing our founders' own call-centre floors over years of operating them, so they are not research and they are not an industry standard, and the exchange below is invented to show the shape rather than transcribed from a recording.

An invented exchange, for shape only

Buried: "Thanks for calling. So, implants. We work with several protocols, and pricing depends on a number of factors, and of course every case is different, so what I can tell you is that..." Leading: "Turnkey, one zone, that is a single figure and I can give it to you now. It covers the consultation, the procedure and the follow-up. Do you want the two-zone range as well?" Same reply speed. Different call.

How to time a demo call

  1. Pick one question you would actually phone about, and know the answer you want before you dial. A vague browse cannot be timed.
  2. Start the stopwatch when you ask it, not when the call connects. The greeting is not the thing you are measuring.
  3. Stop the stopwatch when you have the answer in a form you could act on. A promise to call back is not an answer.
  4. Count the turns it took. The same elapsed time spent in two turns and in nine turns are two different products.
  5. Ask the question again, worded differently, later in the same call. This is where the re-asked question and the confirmation loop show up.
  6. Then, and only then, ask the vendor for their latency figure. You now know what it does and does not predict.

Why rhythm shows up in the numbers, not just the feel

Everything on the second clock is money. A caller who gets a price in seconds is a caller still on the line when the booking question arrives; a caller deep inside a confirmation loop is a caller reaching for the end-call button. That is why the pace of the conversation belongs in the specification of an AI receptionist, next to the accuracy of what it says and the security of where it says it.

The first clock is table stakes, and every serious vendor now clears it. The second clock is where calls are still being lost, and it is the one you can measure yourself, on a single demo call, with the stopwatch already in your pocket.

Frequently asked questions

How fast should an AI voice agent reply?
Fast enough that the pause between turns does not announce the machine. But that figure measures only the gap between turns. What decides whether a call feels fast is how many turns pass before the caller has the answer they phoned for, which is a matter of rhythm rather than latency.
What is speech rhythm in a phone conversation?
The shape of time across the whole call: how long each turn runs, where the answer sits inside it, how many turns pass between a caller giving information and the agent using it, and whether pauses are filled or silent. Latency is one interval inside that shape.
How do I test an AI voice agent's rhythm on a demo call?
Time the call from your question to the answer you could act on, not from the moment you stop speaking, and count the turns it took. Then ask the same question again, worded differently, later in the call: re-asked questions and confirmation loops surface there.
Can an AI voice agent have low latency and still lose the call?
Yes. An agent can begin speaking immediately, then bury the answer at the end of a long turn, read information back twice, or go silent while it looks something up. Every turn is fast and the call is slow, and callers hang up on the call.