The 88% and the 6%
Two numbers from the same year describe the state of AI in customer service better than any vendor deck. McKinsey's 2025 State of AI report found 88% of organisations using AI somewhere in the business, and 6% getting value from it at scale. Everything interesting sits in the gap between those two figures, and almost none of it is a technology problem.
We watch this from the operator's side. In our own operations the founders ran call centres for about ten years, up to 1,500 live operators at peak, across France, Mexico and beyond. The automation projects that went nowhere rarely failed because the software could not hold a conversation. They failed because nobody decided which calls the agent was supposed to take, and nobody agreed in advance how anyone would know whether it took them.
88% of organisations use AI. 6% get value from it at scale. (McKinsey State of AI, 2025)
Value at scale is a strict phrase. It means a line on the P&L moved and stayed moved: fewer overtime hours, a queue that no longer needs a night shift, a booking rate that holds through the ad spike. A pilot that takes a slice of the queue and reports high containment is not that. It is a demo running on production traffic.
Board pressure is not a business case
Gartner reports that 91% of customer-service leaders are under pressure from their leadership to implement AI in 2026. A Gartner survey in 2025 found that 20% of service organisations had actually reduced headcount. Those two numbers sit in the same market, and the distance between them is the most honest description of it currently available.
Pressure from above produces deployments, and deployments get graded on being live. That is how a call centre ends up with a voice agent on the main line, a slide saying AI is in production, and a rota that has not changed by a single shift. The board asked for AI. It got AI. Nobody asked what the line was supposed to stop doing.
Gartner's 20% deserves a fair reading. Not every operation that kept its headcount failed. Some deliberately spent the freed capacity on growth, on calls they used to abandon, or on hours they never staffed, and that is a legitimate outcome and often a better one. But if a project cut no cost, added no coverage and lifted no conversion, then it did not do anything, and a fair amount of that is hiding in the remainder.
The question to ask
Before approving the next phase, ask what changed on the P&L or the rota, not how many conversations the agent handled.
The trajectory is real, the timetable is not yours
The direction of travel is not in dispute. Salesforce's 2025 State of Service puts 30% of service cases as handled by AI today, rising to 50% by 2027. Gartner forecasts that one in ten agent interactions will be automated by AI in 2026, up from roughly 1.6%. Gartner also forecasts that by 2029, 80% of common service issues will be resolved autonomously by AI agents, with an associated cost reduction of around 30%.
Read the 2026 forecast carefully, because it is the one describing the year you are actually in. Gartner's step from roughly 1.6% of interactions to one in ten multiplies the base several times over, and it happens on a base that was tiny to begin with. That is what a market looks like when the easy call types go first: opening hours, order status, appointment changes, the same handful of questions all day. The hard conversations are still in the queue at the end of the decade, which is why Gartner's 2029 forecast covers common service issues rather than everything.
The cost side of the same forecast is blunter. Gartner expects conversational AI to remove $80B of contact-centre agent labour cost in 2026, against a workforce Gartner put at 17M contact-centre agents worldwide in 2022. Both figures are large, and neither one tells you how much of the labour bill survives. Money is coming off the labour line. The labour line is not disappearing.
Where the money actually sits
Two Gartner figures set the unit economics behind every decision in this category. Gartner priced a self-service contact at $1.84 in 2025 against $13.50 for a live assisted one, about seven times the cost. And Gartner put up to 95% of total contact-centre cost as labour in 2022. Almost everything you spend is a person talking.
That is why the enthusiasm is rational rather than faddish. If nearly all the cost is people on calls, then taking a conversation that repeats all day off a person moves real money, and doing it requires no bet on the technology, only a decision about which conversation.
It is also where the arithmetic quietly gets faked. Gartner's $1.84 only holds when the contact is finished. A contact handled halfway that comes back does not replace the expensive number with the cheap one; it adds them together, and it adds a customer who now explains the problem twice and starts the second call irritated. Containment dashboards report the first half of that and never see the second.
$1.84 for a self-service contact, $13.50 for a live assisted one. (Gartner, 2025)
Why deployments stall
After enough of these, the stall pattern is boringly consistent. It is usually one of four things, and none of the four shows up in a vendor evaluation.
- Scope drawn by department instead of by call type.
- Deflection counted as if it were resolution.
- No exit path to a person, so every limit becomes a fight with the caller.
- Nobody listening to the calls, so the failure is invisible until it shows up in churn.
Scope by department is the most common failure and the most expensive. Support is not a scope, it is an org chart. It holds the routine questions an agent could answer on day one together with the conversations that need a person with authority to make an exception, and it hides both behind one word. A neuro-agent scoped that way has to be competent at everything before it is allowed to be useful at anything.
Deflection is the metric that flatters everyone who reads it. It counts calls that did not reach a person, which includes the calls that were solved and the calls where the customer gave up. Resolution counts the ones where the reason for calling stopped existing. Those two numbers can differ by a wide margin, and only one of them predicts whether the customer calls back on Monday.
You also cannot tell them apart from a dashboard. In our own operations under 5% of recorded calls were ever reviewed under manual QA, which means the normal way of checking a rollout is a sample too small to catch a systematic failure. If the agent mishandles one intent in a way that sends every one of those callers back to the queue, a sample that size will find it eventually, well after the callers already have.
What the 6% do differently
Nothing exotic. Three habits, and all three are decisions made before anything gets configured.
- Scope to one call type that repeats, narrow enough that it feels underambitious.
- Define resolution in the customer's terms before launch, and audit it by listening to calls.
- Keep a human exit path and announce it in the greeting.
Start with the call type that repeats, and pick it narrow: one intent, high volume, low variance, a clear finish line. Narrow scope is what makes measurement possible. When the agent handles one thing, the resolution rate for that thing is legible and so is the failure. Wide scope averages the failure away until it is invisible.
Then measure resolution and check it by listening. Choose the definition before launch, in the customer's language rather than the system's, and verify it against real conversations instead of a containment counter. This is the one habit that separates a rollout improving month over month from one that plateaus in its first weeks and gets quietly redescribed as a pilot.
And keep the exit path. Callers hear it as reassurance. Operationally, it is the thing that makes aggressive automation safe to attempt. SurveyMonkey found in 2025 that 8% of consumers prefer AI over a human, and McKinsey found in 2024 that 71% of Gen Z pick up the phone when self-service fails. A route to a person that the caller can hear means an agent hitting its limit hands the call over instead of arguing, so the worst outcome of a bad automation decision is a transfer, not a lost customer.
The gap McKinsey measured will close, and not mainly because the models get better. It will close for the operations that treat a rollout as an operational change with a number attached to it, and it will stay open for the ones that treat it as a purchase. Board pressure sets the deadline. It does not do the scoping.