McKinsey's 2025 State of AI report put 88% of organisations as users of AI and 6% as getting value from it at scale. A good part of that gap is a definition problem. Somebody signed off on a thing described as an AI agent, and what arrived was a chat widget with a decision tree behind it.
So the definition first, because it is duller than the marketing and much better at predicting what you will get. An AI agent is software you give a goal, a set of tools and a set of rules. It works out the steps itself. A script does the opposite: somebody wrote every branch in advance, and whatever falls outside the branches falls off the end of the call.
We build these for the phone, which is the least forgiving surface for them, and we call ours neuro-agents, so read this as a partisan article. It is still a plain description — what the word covers, what the parts are, what the thing does well, and where it is worse than the person you were going to hire.
Where this comes from
Ten years of running call centres, up to fifteen hundred live operators at peak, across France, Mexico and beyond. That history is where the operating figures below come from — how little of a call archive anybody ever listens to, and what our own reply latency actually measures. Everything attributed to a named publisher is cited in the sentence that uses it. None of it is research and none of it is an industry standard.
What an AI agent actually is
Four things have to be true before the word means anything.
- A goal it holds. Not a step to execute but an outcome to reach: get the appointment booked, find out whether this lead has a budget, confirm tomorrow's delivery window.
- Tools it can use. Your CRM, your calendar, your price list, a warm transfer to a colleague. Without tools an agent can describe your business and cannot act in it.
- Rules it cannot break. What must never be promised, what must always be disclosed, the point at which a human takes the call. This is the part buyers skim and later wish they had not.
- The conversation so far. It remembers what the caller said earlier in the call and answers in light of it.
Remove the goal and you have an autoinformer: it plays its message and hangs up. Remove the tools and you have a chatbot that explains your booking system to someone who wanted a booking. Remove the rules and you have the thing that ends up quoted in a newspaper.
The difference is easiest to hear on an example, and this one is invented rather than transcribed. A caller asks whether you deliver to a district you do not cover. A script has a branch for delivery and a branch for postcodes, and this question sits between them, so it fires the branch nearest the keyword and says something confident and useless. An agent holding the goal — take the order if we can serve them — checks the coverage list, says no, and offers the nearest pickup point instead. Same input. Different machine.
An AI agent is not a chatbot
The two get sold under one heading and they break in different places.
A chatbot maps the conversation in advance. Someone drew the tree; each incoming message is matched to the nearest branch. That holds while people ask the expected questions in the expected words, and it comes apart the moment they don't. Five9's 2024 consumer survey found 56% of consumers often frustrated by service chatbots, and the Observatoire des Services Clients (Ipsos BVA, 2025) scored chatbot satisfaction in France at 50% — the lowest-rated channel in the study.
An agent has no tree. It has the goal, the rules and whatever it can look up, and it composes the answer each time. Which is the upside and the exposure in one sentence: it handles the question nobody anticipated, and it will also compose a wrong answer if your knowledge base contradicts itself.
Then there is the difference that decides most projects and gets mentioned least. Chat gives you time. A person types, waits, looks at something else, comes back long afterwards, and nobody minds. A phone call gives you about a second. Silence on the line reads as a dropped connection — the caller says hello and starts the sentence again. Nearly everything hard about a voice agent comes out of that one budget. If the distinction you actually need is against an IVR or a callbot, we drew that line separately in callbot, voicebot or IVR.
What businesses put them on the phone to do
Across the accounts we run, the work falls into five jobs. Each has a number attached to it at the end of the month, which is the only honest reason to automate any of them.
- Inbound calls nobody is answering. The 411 Locals study in 2016 put 62% of calls to small businesses as unanswered; ContactBabel's 2024 UK data has the average caller waiting 116 seconds before a human voice, and 8.4% hanging up before anyone answers. The agent picks up on the first ring, in the middle of the night, on every line at once.
- Working a list. New enquiries, dormant customers, a database somebody bought and never worked. Harvard Business Review, on MIT's lead-response data, put a company calling a new lead within five minutes at 21x more likely to qualify it than one calling at thirty minutes — and found that 1% of the B2B companies studied actually respond inside five minutes. That is a staffing-hours problem, and an agent is a staffing-hours answer.
- Confirmations and reschedules. Tomorrow's appointments, tomorrow's deliveries. Psychiatric Services (APA, 2018) recorded a 39% no-show rate when a reminder call went unanswered against 3% when a live reminder conversation actually happened.
- The whole sale. First contact through to a placed order, objections included. This is the ambitious job and the one that needs the most of your own recordings before it is any good.
- Reading the calls your people already make. In our own operations, under 5% of recorded calls are ever reviewed by hand. Locator listens to 100% of them against 30+ parameters and files what was said back onto the card.
What happens inside one call
Three things run in series on every turn, and the whole chain has to finish before the caller notices a gap.
- Speech to text. The caller's audio becomes words while they are still speaking, together with a judgement about whether they have finished the thought.
- The decision. A language model reads those words with the conversation so far, your knowledge base and your rules, and writes the reply.
- Text to speech. The reply becomes audio in the agent's voice.
Serial, not parallel: each link waits for the one before it. Past roughly a second of silence callers start filling it themselves, repeating the question or asking whether anyone is there. Ours replies in about 0.3 seconds, with recognition and synthesis running on our own servers — the detail is on the technology page. Most engineering time on an AI voice agent goes into that chain, and it is why a working chat agent does not become a working phone agent by changing the output device.
The other half of the work is teaching it when to say nothing. A person pausing mid-sentence to think has not finished talking, and a script has no way to tell that silence from the other kind. So the familiar bot cuts in, and after the second interruption the caller stops treating it as a conversation. That failure has its own article — why bots interrupt — because it is the most common reason people call an agent robotic when the voice itself is fine.
What it does better than a person, and where it loses
The strengths are the ones you would guess, and they are worth stating precisely. It does not tire, so the last call of the day is delivered like the first. It never forgets a callback, because the callback is a row rather than a memory. It works nights, weekends and the stretch when the whole team is at lunch. And it scales sideways in a way a floor cannot: a hundred simultaneous calls are answered the same way as one, up to the capacity agreed for your line.
Now the part that stays off the slide. It reads fine emotion worse than a competent human — a caller who is technically satisfied and quietly furious gets handled correctly rather than well. It is cautious off-script, which is the right instinct and means an unusual request turns into a handover rather than a resolution. It cannot run a complicated deal with several signatories and a legal review. And a knowledge base that disagrees with itself will produce one of the disagreements, stated confidently, on a live call.
None of that is an argument against putting one on the phone. It is an argument against buying one to replace a floor. Gartner's 2025 survey found 20% of service organisations actually cut headcount after deploying AI. What happened in the other organisations is not in that survey, and we are not going to invent a figure for it; on our own accounts the work moved rather than disappeared. Five9's 2024 consumer survey still has 75% of consumers preferring to talk to a real human, and SurveyMonkey put the share who prefer AI over a human at 8% in 2025. Design the handover properly and neither number is your problem.
A script knows every question somebody thought of in advance. An agent is what you put on the line for the ones nobody did.
What it costs
Pricing is per minute of conversation, and ours is published rather than quoted, because a quote for a call minute tends to be a quote for how much you look like you can pay.
- Locator, speech analytics on the calls your own team makes: EUR 0.10 a minute, the same nominal figure in USD.
- Bene Hotline, inbound reception and routing: EUR 0.20.
- Bene Qualifier, lead qualification on a list: EUR 0.25.
- Bene Sales, first contact through to a placed order: EUR 0.30.
All excluding VAT. Volume discounts reach -25% on the agents and -62.5% on Locator, and a no-commitment trial carries a +20% premium over the base rate — you pay a little more for the right to stop. The calculator prices a real volume before you talk to anybody here.
The comparison that actually decides the question is not one vendor against another. Gartner priced a self-service contact at $1.84 against $13.50 for a live assisted one in 2025, roughly seven to one, and puts labour at up to 95% of contact-centre cost (2022). Those two figures are the business case. They are also why the number to argue about is the share of calls that finish without a person, not the price of the minute — we took that arithmetic apart in cost per minute.
What you need before you start
What we ask for is less than people expect and more specific.
- Recordings. Twenty or thirty real calls, the good ones and the bad ones. The agent is built against how your customers actually ask, not against how your website describes what you sell.
- A price list and a coverage list that agree with themselves. Every contradiction in there is a confident wrong answer waiting for a caller.
- The rules. What may never be promised, what must always be said, and the point at which a human takes over.
- One person who can decide. Not a committee that meets on Thursdays.
Launch takes three days, because the agent is configured against those recordings and your existing pipeline rather than built from nothing. The week after launch is calibration, and anyone who tells you the first week is production has not run one.