AI Agent Technology · 17 August 2026 · 13 min read

The exit path: when an AI voice agent should hand the call to a human

Most deployments script the automated part of the call in detail and leave the escalation to chance. The handoff is the part callers judge you on, and it decides whether the automation gets trusted at all.

WHO CALLERS ASK FORConsumers who still prefer a real human75%Consumers who prefer AI to a human8%Most callers still want a person, so the route to one is part of the product.

The handoff is part of the product, not the failure case

There is a version of the voice-agent pitch in which transferring to a human is an embarrassment: the moment the machine admits it could not cope. It is a bad way to think about it and it produces bad deployments. Everyone building one of these already knows which conversations the agent will not finish. Almost nobody writes those conversations down with the care they give to the ones it will.

The preference data is not close. A Five9 consumer survey in 2024 found that 75% of consumers still prefer talking to a real human, while SurveyMonkey in 2025 put the share who prefer AI over a human at 8%. Those two figures do not describe people who hate automation; the same people use self-service constantly and without complaint. They describe what a caller wants the moment a call stops being routine. When that moment arrives, the exit path is not a rare branch of the script. It is the thing the caller was already looking for.

75%
consumers who still prefer talking to a real human (Five9 consumer survey, 2024)
8%
consumers who prefer AI over a human (SurveyMonkey, 2025)

The corroboration runs the same direction across unrelated sources. Five9 also found that 56% of consumers are often frustrated by service chatbots. In France, the Ipsos BVA Observatoire des Services Clients rated chatbot satisfaction at 50% in 2025, the lowest-rated channel in the study. And McKinsey found in 2024 that 71% of Gen Z pick up the phone when self-service fails, which is worth sitting with: the cohort most often assumed to have abandoned the telephone is the one that reaches for it the moment an automated path breaks.

The ContactBabel series makes the point most plainly of all: phone preference for urgent, complex issues rose from 28% to 47% between 2018 and 2024, on ContactBabel's own measurement. The phone is not losing ground to automation. It is specialising into exactly the calls that automation cannot finish, which means the volume arriving at your exit path is made of your hardest and most valuable conversations, not your easiest ones.

The reframe

A transfer is not the moment the automation failed. It is the moment the caller finally gets the thing most of them said they wanted, and the way it is executed is what they will remember about the entire call.

Three exits, and only one of them is a defect

Lumping every transfer together under "the bot could not do it" hides the fact that they are three different events with three different fixes. A deployment that reports one aggregate transfer rate cannot tell you which of them it is looking at.

The reason to separate them is that the first two are capacity questions and the third is an engineering question. Reported as a single number, a rise gets treated as a quality problem, and the usual response is to make the agent more reluctant to transfer. That is the opposite of the fix in two cases out of three, and it converts a working exit into a trap.

What a clumsy exit actually costs

The cost is almost never the transfer itself. It is the wait and the repetition on the other side of it, and both of those are inherited from the switchboard the automation was supposed to improve on.

ContactBabel put the average wait before a caller hears a human voice at 116 seconds in UK contact centres in 2024, with 8.4% of inbound callers hanging up before anyone answers at all. A Velaro survey back in 2012 found that around 60% of callers will not hold longer than one minute. Route someone who has already spent ninety seconds explaining themselves to an agent into a two-minute queue, and you have built a worse experience than the one you replaced, using better technology to do it.

Inbound callers who hang up before anyone answers8%
8.4%, ContactBabel, 2024
Callers who will not hold longer than one minute60%
around 60%, Velaro survey, 2012
Dissatisfied customers who never complain and simply switch56%
Zendesk CX Trends, 2026

The second cost is the one that goes unmanaged, because it is invisible by construction. Zendesk CX Trends 2026 found that 56% of dissatisfied customers never complain and simply switch, and that more than half of customers switch after a single bad experience. A botched handoff does not generate a complaint you can count. It generates a customer you never hear from again and a transfer log that still looks perfectly healthy.

What the human on the other end has to receive

A handoff that works is not a transfer. It is a transfer plus a briefing, and the briefing is what separates it from a cold restart. Four things have to travel with the call.

01
Why the call is being transferred

Which of the three exits this is, in one line, so the person picking up knows immediately whether they are taking a routine specialist question or a rescue. The two want completely different opening sentences.

02
What the caller already said

Name, account or booking reference, and the actual problem in the caller's own words rather than a category code. The most reliable way to make an already irritated customer feel worse is to make them say all of it a second time.

03
What the agent already promised

Anything the automated part of the call committed to, whether that is a callback window, a price band or a slot held, travels with the call. Otherwise the human contradicts it inside the first sentence and the company looks disorganised rather than automated.

04
Where to reach them if the line drops

A transfer is the most fragile moment in any call. A number and a one-line note survive a dropped line. The conversation frequently does not, and the caller is not going to be the one who tries again.

This is not new machinery. Our qualifier neuro-agent already writes exactly this object, a lead card, into the CRM on every call it handles. An escalation is the same card with a different urgency attached to it, which is why the handoff format is worth deciding once and reusing rather than inventing per campaign.

Design the exit before you write the script

The order matters more than it looks. Teams that write the happy path first end up bolting the escalation on at the end, at which point it inherits whatever routing already existed. Four decisions, taken before any dialogue gets written.

  1. Decide what the agent may never handle. Write that list before the happy path, not after it. It is a commercial and legal decision about liability and tone, and it should not be made by whoever is editing prompts on a Thursday.
  2. Decide who receives it, hour by hour. An exit path with nobody behind it at seven in the evening is not an exit path, it is a dead end with a polite voice in front of it. Map the receiving rota against your actual call distribution, not your office hours.
  3. Decide what happens when nobody is there. This is the decision most often skipped, and the default it falls back to is voicemail.
  4. Decide the words. "I am putting you through to a colleague now, they will already have everything you have told me" produces a different call from "transferring you". The first sets an expectation the briefing then meets.

The voicemail default deserves naming as its own decision, because the numbers on the far side of it are poor. CallRail put the share of callers who leave a voicemail at 42% in 2025, and an industry estimate attributed to BIA/Kelsey, an estimate rather than a measured study, holds that 85% of people who reach voicemail never call back. An out-of-hours escalation is better served by taking an explicit callback commitment with a stated window, which the agent can do, than by handing the caller a recording and hoping.

The cheap test

Ask a colleague to call your own line, say the words "I would like to speak to someone please", and time everything that happens next. Most teams discover the answer to their escalation design that afternoon.

How you know the exit path is working

Three numbers, and all three have to be measured on every call rather than on a sample.

Sampling will not surface any of them. In our own operations the share of recorded calls that ever got reviewed under manual QA stayed under 5%, and review meant a handful of calls per agent per month, chosen by whoever had the time that week. Escalations are by definition the calls least likely to appear in that handful, because they are the exceptions and the sample is drawn from the ordinary. Scoring 100% of calls is what turns the exit path from an anecdote somebody remembers into a number somebody owns.

Recorded calls ever reviewed under manual QA5%
under 5%, in our own operations
Calls scored automatically100%
our operating practice, not a research finding

What this means for the business

A voice agent with a designed exit path is a different commercial proposition from one without. It can be given the ordinary volume without anyone having to pretend it will handle everything, because the boundary is written down and the route across it is short. That is what makes the deployment survivable internally: the operators who feared being replaced can see precisely which calls are still theirs, and they are the interesting ones.

It also changes what a rising transfer rate means. Split three ways, it becomes readable. Out-of-scope volume going up is a product signal about what customers are actually calling about. Explicit requests going up is a trust signal about how the agent is being received. Only the third line is a defect, and it is the one you can fix on Monday by changing a script line rather than by retraining a person.

The uncomfortable version of the same point: if you cannot say today what happens when a caller asks your automated line for a human at seven in the evening, that answer already exists. Someone is receiving it, or nobody is. It was simply decided by default rather than on purpose, and the customers who found it are not going to write in about it.

Frequently asked questions

When should an AI voice agent transfer a call to a human?
In three situations, and they are worth counting separately. First, when the call is about something the agent was deliberately never given, such as a dispute, a bespoke quote or a complaint about a person. Second, whenever the caller clearly asks for a human. Third, when the conversation itself has gone wrong: repeated mishearing, looping, or a caller who is angry, distressed or describing something urgent. Only the third is a defect; the first two are the design working.
Should the agent transfer as soon as the caller asks for a human?
Yes, on the first clear request, and without re-offering self-service first. A Five9 consumer survey in 2024 found 75% of consumers still prefer talking to a real human, against 8% who prefer AI over a human in SurveyMonkey's 2025 figure. Arguing with that preference costs more than the call it saves. The practical work is recognising the request in all the ways people phrase it, including impatient and rude phrasings that never appear in a script written at a desk.
What information should be passed to the human agent during a transfer?
Four things: why the call is being transferred, what the caller has already said in their own words including any reference number, anything the automated part of the call already committed to such as a callback window or a held slot, and a number to reach the caller on if the line drops. Without the third item the human contradicts a promise the caller has already been given, which is worse than no automation at all.
Does a high transfer rate mean the AI voice agent is not working?
Not on its own, and treating it that way tends to make the deployment worse. Split the rate into out-of-scope transfers, explicit caller requests and conversations that went wrong. The first is a product signal, the second is a trust signal, and only the third is a fault. Responding to an aggregate rise by making the agent more reluctant to transfer is the wrong fix in two of those three cases.
What should happen if no human is available to take the transfer?
Decide it deliberately rather than letting it fall back to voicemail. CallRail put the share of callers who leave a voicemail at 42% in 2025, and an industry estimate attributed to BIA/Kelsey, an estimate rather than a measured study, holds that 85% of those who reach voicemail never call back. An explicit callback commitment with a stated window, taken by the agent and written into the CRM, holds far more of the demand than a recording does.
How do you measure whether the handoff is actually working?
Three numbers, measured on every call rather than on a sample: transfer rate split by the three exit types, time from the caller's first clear request to a human voice, and how often the human has to ask for something the caller already gave the agent. That last one predicts the satisfaction score and is almost never logged. In our own operations the share of recordings ever reviewed under manual QA stayed under 5%, and escalations are the calls least likely to land in that sample, which is why the exits have to be scored automatically on all of them.