Barge-in is the moment the caller starts talking while the agent is still talking. Every vendor demo includes one, because it is easy to stage and it always looks good: the agent stops, the caller finishes, the conversation carries on. What the demo cannot show you is the part that decides whether the call actually worked.
Stopping is the easy half. Any system with voice-activity detection can stop. The hard half is what the agent does next, and answering that correctly requires a number most platforms never record.
The three questions an interruption asks
When a caller cuts in, the agent has to resolve three things at once, and they are not equally difficult.
- Has the caller actually started speaking, or was that a cough, a door, or a television in the background?
- How much of what I was saying did they hear before they started?
- Given that, what do I need to say again before I move on?
The first question is solved. The third is a scripting decision. The second is the one that separates platforms, because it is the only one that requires the system to have been keeping count.
Stopping when the caller speaks is table stakes. Knowing what they missed is the part that changes the next sentence.
The field to look for: heard share
Speech is generated and delivered progressively. At the instant the caller interrupts, some measurable portion of the reply has been spoken and the rest has not. A platform that records that boundary can express it as a share of the intended response — how much landed, how much was lost.
Here is one of ours, from a logged call. The agent set out to deliver a response of 800 characters. The caller cut in after 91 of them. That is roughly a tenth of the answer, which means the other nine tenths — including, in that instance, the part that answered the question — never reached the caller at all. An agent that treats the interruption as an acknowledgement and advances to the next step has just built the rest of the call on something the caller never heard.
A second logged interruption, from a technical test against a demo agent of our own, sat closer to the middle: the caller had heard 76 of the 159 characters the agent was part-way through saying, which is 47% of the reply. The decisive detail sat in the unheard half. Because the platform had the number, the agent cancelled its own move to the next stage of the scenario, restated the substance, and only then went back to its question.
The zero case, and why it matters most
Three further interruptions in that same test came earlier in the turn: 0% heard, each time. Nothing of the reply had landed. That is the cleanest signal a log can give you, and it is also the one that most obviously should change the agent's behaviour — there is no point rephrasing or summarising something the caller has not heard a word of. The correct move is to acknowledge, deliver the point in full, and only then continue.
An agent without the measurement cannot distinguish that case from the one above. It has two options, and both are wrong some of the time: repeat everything, which makes it sound broken, or repeat nothing, which makes it sound like it was not listening. The number is what turns a guess into a decision.
The question to put to a vendor
Not “does it handle barge-in”. Ask: does the platform record what proportion of each reply the caller heard before interrupting, is that figure in the call log where I can read it, and does the agent use it to decide what to repeat? The first answer is always yes. The other two are where vendors separate.
Reading a log you have been given
If you are evaluating a platform, ask for a raw call log rather than a summary dashboard, and look for three things in it.
- An interruption event that is timestamped against the agent's own output, not just against the call clock.
- A heard-share or characters-delivered figure attached to that event — a percentage or a count, either is fine, absent is not.
- Evidence in the next turn that the agent did something different because of it. A log that records the number and then carries on identically is instrumentation without a decision behind it.
That third point is the one worth being strict about. Recording a measurement is cheap. Acting on it changes what the agent says, and it is visible in the transcript when it does.
A footnote on reading confidence scores
One more line from the same logs, offered without a conclusion attached because we do not have one. The platform's own speaker analysis first classified our human tester as “robot” with 75% confidence, then on re-analysis as “human” with 97%. We are not sure what to make of a bot-detector that briefly suspects a person, except that it is a useful reminder to read a confidence score as a confidence score and not as a verdict — including the ones a vendor puts in front of you.
Why this is worth your time
Interruptions are not an edge case. Callers interrupt when they are impatient, when they already know the answer, and when they are trying to get to a person — which is to say, on the calls that matter most. An agent that mishandles them fails quietly, in a transcript that reads as a polite conversation that went nowhere.