The handoff is part of the product, not the failure case
There is a version of the voice-agent pitch in which transferring to a human is an embarrassment: the moment the machine admits it could not cope. It is a bad way to think about it and it produces bad deployments. Everyone building one of these already knows which conversations the agent will not finish. Almost nobody writes those conversations down with the care they give to the ones it will.
The preference data is not close. A Five9 consumer survey in 2024 found that 75% of consumers still prefer talking to a real human, while SurveyMonkey in 2025 put the share who prefer AI over a human at 8%. Those two figures do not describe people who hate automation; the same people use self-service constantly and without complaint. They describe what a caller wants the moment a call stops being routine. When that moment arrives, the exit path is not a rare branch of the script. It is the thing the caller was already looking for.
The corroboration runs the same direction across unrelated sources. Five9 also found that 56% of consumers are often frustrated by service chatbots. In France, the Ipsos BVA Observatoire des Services Clients rated chatbot satisfaction at 50% in 2025, the lowest-rated channel in the study. And McKinsey found in 2024 that 71% of Gen Z pick up the phone when self-service fails, which is worth sitting with: the cohort most often assumed to have abandoned the telephone is the one that reaches for it the moment an automated path breaks.
The ContactBabel series makes the point most plainly of all: phone preference for urgent, complex issues rose from 28% to 47% between 2018 and 2024, on ContactBabel's own measurement. The phone is not losing ground to automation. It is specialising into exactly the calls that automation cannot finish, which means the volume arriving at your exit path is made of your hardest and most valuable conversations, not your easiest ones.
The reframe
A transfer is not the moment the automation failed. It is the moment the caller finally gets the thing most of them said they wanted, and the way it is executed is what they will remember about the entire call.
Three exits, and only one of them is a defect
Lumping every transfer together under "the bot could not do it" hides the fact that they are three different events with three different fixes. A deployment that reports one aggregate transfer rate cannot tell you which of them it is looking at.
- Out of scope, by design. The call is about something the agent was deliberately never given: a dispute, a bespoke quote, a contract term, a complaint about a person. Nothing failed here. The routing worked exactly as specified. These should be the bulk of your transfers and their share should be stable week to week.
- The caller asked. Someone says they want a person. This is not a negotiation and it is not the place to re-offer self-service. The correct behaviour is to comply on the first clear request, in all of the ways people actually phrase that, including the impatient and the rude ones.
- The conversation went wrong. The agent misheard twice, looped on the same question, or the caller is angry, distressed or describing something urgent. This is the only one of the three that is a defect, and it is the only one whose number should be falling week over week.
The reason to separate them is that the first two are capacity questions and the third is an engineering question. Reported as a single number, a rise gets treated as a quality problem, and the usual response is to make the agent more reluctant to transfer. That is the opposite of the fix in two cases out of three, and it converts a working exit into a trap.
What a clumsy exit actually costs
The cost is almost never the transfer itself. It is the wait and the repetition on the other side of it, and both of those are inherited from the switchboard the automation was supposed to improve on.
ContactBabel put the average wait before a caller hears a human voice at 116 seconds in UK contact centres in 2024, with 8.4% of inbound callers hanging up before anyone answers at all. A Velaro survey back in 2012 found that around 60% of callers will not hold longer than one minute. Route someone who has already spent ninety seconds explaining themselves to an agent into a two-minute queue, and you have built a worse experience than the one you replaced, using better technology to do it.
The second cost is the one that goes unmanaged, because it is invisible by construction. Zendesk CX Trends 2026 found that 56% of dissatisfied customers never complain and simply switch, and that more than half of customers switch after a single bad experience. A botched handoff does not generate a complaint you can count. It generates a customer you never hear from again and a transfer log that still looks perfectly healthy.
What the human on the other end has to receive
A handoff that works is not a transfer. It is a transfer plus a briefing, and the briefing is what separates it from a cold restart. Four things have to travel with the call.
Which of the three exits this is, in one line, so the person picking up knows immediately whether they are taking a routine specialist question or a rescue. The two want completely different opening sentences.
Name, account or booking reference, and the actual problem in the caller's own words rather than a category code. The most reliable way to make an already irritated customer feel worse is to make them say all of it a second time.
Anything the automated part of the call committed to, whether that is a callback window, a price band or a slot held, travels with the call. Otherwise the human contradicts it inside the first sentence and the company looks disorganised rather than automated.
A transfer is the most fragile moment in any call. A number and a one-line note survive a dropped line. The conversation frequently does not, and the caller is not going to be the one who tries again.
This is not new machinery. Our qualifier neuro-agent already writes exactly this object, a lead card, into the CRM on every call it handles. An escalation is the same card with a different urgency attached to it, which is why the handoff format is worth deciding once and reusing rather than inventing per campaign.
Design the exit before you write the script
The order matters more than it looks. Teams that write the happy path first end up bolting the escalation on at the end, at which point it inherits whatever routing already existed. Four decisions, taken before any dialogue gets written.
- Decide what the agent may never handle. Write that list before the happy path, not after it. It is a commercial and legal decision about liability and tone, and it should not be made by whoever is editing prompts on a Thursday.
- Decide who receives it, hour by hour. An exit path with nobody behind it at seven in the evening is not an exit path, it is a dead end with a polite voice in front of it. Map the receiving rota against your actual call distribution, not your office hours.
- Decide what happens when nobody is there. This is the decision most often skipped, and the default it falls back to is voicemail.
- Decide the words. "I am putting you through to a colleague now, they will already have everything you have told me" produces a different call from "transferring you". The first sets an expectation the briefing then meets.
The voicemail default deserves naming as its own decision, because the numbers on the far side of it are poor. CallRail put the share of callers who leave a voicemail at 42% in 2025, and an industry estimate attributed to BIA/Kelsey, an estimate rather than a measured study, holds that 85% of people who reach voicemail never call back. An out-of-hours escalation is better served by taking an explicit callback commitment with a stated window, which the agent can do, than by handing the caller a recording and hoping.
The cheap test
Ask a colleague to call your own line, say the words "I would like to speak to someone please", and time everything that happens next. Most teams discover the answer to their escalation design that afternoon.
How you know the exit path is working
Three numbers, and all three have to be measured on every call rather than on a sample.
- Transfer rate split by the three exits above. One aggregate figure tells you nothing you can act on, and it moves for reasons that cancel each other out.
- Time from the request to a human voice, measured from the caller's first clear request rather than from the moment the system began the transfer. The gap between those two timestamps is usually where the problem lives.
- Repetition rate: how often the human has to ask for something the caller already gave the agent. This is the number that predicts the satisfaction score, and it is the one almost nobody logs.
Sampling will not surface any of them. In our own operations the share of recorded calls that ever got reviewed under manual QA stayed under 5%, and review meant a handful of calls per agent per month, chosen by whoever had the time that week. Escalations are by definition the calls least likely to appear in that handful, because they are the exceptions and the sample is drawn from the ordinary. Scoring 100% of calls is what turns the exit path from an anecdote somebody remembers into a number somebody owns.
What this means for the business
A voice agent with a designed exit path is a different commercial proposition from one without. It can be given the ordinary volume without anyone having to pretend it will handle everything, because the boundary is written down and the route across it is short. That is what makes the deployment survivable internally: the operators who feared being replaced can see precisely which calls are still theirs, and they are the interesting ones.
It also changes what a rising transfer rate means. Split three ways, it becomes readable. Out-of-scope volume going up is a product signal about what customers are actually calling about. Explicit requests going up is a trust signal about how the agent is being received. Only the third line is a defect, and it is the one you can fix on Monday by changing a script line rather than by retraining a person.
The uncomfortable version of the same point: if you cannot say today what happens when a caller asks your automated line for a human at seven in the evening, that answer already exists. Someone is receiving it, or nobody is. It was simply decided by default rather than on purpose, and the customers who found it are not going to write in about it.