Most claims about voice agents are made against a demo. This one was made against 436 people doing the same job on the same traffic in the same month, which is a harder audience and a more useful one.
The result is not the one a vendor would pick. The agent did not beat the floor. It came second, narrowly, and second turned out to be the interesting answer.
How we measured
One anonymised advertiser selling impulse goods off a television shopping channel in Moscow, an account of ours, with four offers running. Inbound calls were split live between the client's own sales floor and a neuro-agent on the same scripts and the same offers. June was calibration, July was production. The metric is approval rate: orders confirmed divided by calls handled. These are our own operational numbers, not a study, and no publisher stands behind them. No cost or price from that account appears here, because labour costs do not travel between markets and a ratio built out of them would be a claim about our market pretending to be a claim about yours.
What approval rate is, and why it is a brutal metric
Approval rate is confirmed orders over calls handled. Not leads, not interest, not a callback agreed: an order the customer actually confirmed on that call. It is the least forgiving number in direct-response telesales because there is nowhere for a bad conversation to hide. A call either ends in an order or it does not.
It also has a hard floor underneath it. Below a certain approval rate the media spend that generated the call does not pay for itself, and in this niche that threshold sits in the low thirties. That is the number to keep in mind for everything below: not whether the agent was good, but whether it cleared the line where the account stops losing money.
Month one to month two
June was calibration and it looked like calibration. The agent handled real traffic, got things wrong, and every wrong thing went back into the scripts and the objection handling that week.
July was production. On the lead offer, approval went from 22% to 31%. Across the four offers the month-over-month gain ran from 9 to 14 points, which is a suspiciously tidy pattern until you realise what it actually measures: not the model getting cleverer, but a month of being told what customers on this particular product actually ask.
The gain between month one and month two was not intelligence. It was the transcripts.
By July the agent was landing between 30% and 34% depending on the offer. The client's own floor, on the same traffic in the same weeks, was at 33%.
Read that carefully
The agent lost. On the average it lost by a point or two, and on the strongest offers it drew. Anybody selling you a voice agent on the promise that it converts better than your people is selling you something we did not observe.
Where it landed among actual operators
A floor average hides the thing that matters, because a floor is not a homogeneous machine. Split the same 436 people by how well they actually sell and the picture changes:
- Masters: 34% to 43% approval. Only the very top of the agent's range reaches the bottom of this band, and nothing automated is matching the middle of it this year.
- Mid-level: 31% to 40%. This is the bulk of the floor and the band the agent overlapped with.
- Weak operators: 23% to 28%. The agent beat this band on every offer.
- Novices: 11% to 18%. Not a comparison, a gap.
- The agent: 29% to 34%, sitting on the boundary between mid-level and weak depending on how mature the offer was.
So the honest summary is that the agent performs like a competent mid-tier operator who never has a bad shift. That is a much less exciting sentence than the one on most vendor slides and a much more useful one for planning, because a floor is mostly not masters.
The part the conversion number does not show
While that was happening, the agent's volume moved in a way no floor can copy. On one offer it went from 417 calls handled in June to 2,463 in July. On another, from 569 to 2,338. By the end of July it was taking roughly half of all traffic on the maturest offers.
That is the actual mechanism, and it is not about being better at selling. Traffic off a television channel does not arrive evenly. It arrives in a spike, when the segment airs, and a floor sized for the spike is idle the rest of the day while a floor sized for the average drops the spike on the ground.
The floor does not lose the spike because it sells badly. It loses the spike because it is finite.
Where a live operator's month actually goes
One month of the client's own floor time, as their own reporting had it:
- 50% idle, waiting for a call.
- 31% actually in conversation with a customer.
- 5% on breaks.
- 5% in training.
- 4% rest.
- 4% after-call work.
Roughly a third of paid time is spent talking to customers. That is not a criticism of anybody on that floor; it is the arithmetic of staffing for a peak you cannot predict. Gartner put the labour share of contact-centre cost at up to 95% in 2022, which means the half of the day nobody is talking is very close to the whole cost structure.
Why we are not printing a cost ratio
We measured one on this account and we are deliberately leaving it out. A cost ratio between an agent and an operator is built out of local wages, and quoting a Moscow ratio to a reader in Paris or Mexico City would be inventing their business case for them. Gartner's own 2025 figures put a self-service contact at 1.84 dollars against 13.50 dollars for a live assisted one, roughly seven to one; use that as the shape and your own payroll for the number.
What this does not prove
One advertiser, four offers, two months, one market. It shows what happened there, and it does not show what happens on a considered purchase, on a B2B line, or on anything where the customer needs a week to decide. Impulse goods off a television channel are close to the easiest possible case for automation: short call, narrow question set, decision made on the call.
It also does not show headcount coming out. It did not, and it usually does not: a Gartner survey in 2025 found only 20% of service organisations actually cut headcount after deploying AI. What happened here is that a spike stopped being dropped, which shows up in revenue rather than in payroll.
What is worth copying from this
- Split live traffic, do not run a pilot alongside reality. The only comparison worth anything is the same offers, the same hours, the same customers.
- Give it a calibration month and expect it to be mediocre. Nine to fourteen points of the result came out of that month, and skipping it means shipping the version that had not heard the objections yet.
- Measure the metric the account is actually judged on. Approval rate is unpleasant precisely because it cannot be gamed with a warmer tone.
- Point it at the peak first. The overlap band with your mid-tier operators is where it earns, and the peak is where your floor is losing calls it will never see.
- Compare against your tiers, not your average. If it sits above your weak band, that is a real result, and your average was never a person anyway.
The claim we will stand behind from these two months is narrow and worth more than a bigger one: on short transactional calls, an agent sells like a solid mid-tier operator, all day, at whatever volume arrives.