Most claims about voice agents are made against a demo: a scripted caller, a happy path, a recording picked afterwards. A demo shows that a neuro-agent can hold a conversation. The owner of a sales line needs a second answer: whether it sells as well as the people already on the phones.
Only a comparison on live traffic gives that answer. It is harder to set up than a demo and can embarrass either side, which is exactly why it is worth running.
A method for your own test
This article sets out how to run the comparison and how to read it. It reports no conversion rate, headcount or cost from any account of ours, because a number from one line, one product and one market would tell you little about yours. The only figures below are Gartner's, each with its year.
What approval rate is, and why it is a brutal metric
Approval rate is confirmed orders over calls handled. Not leads, not interest, not a callback agreed: an order the customer actually confirmed on that call. It is the least forgiving number in direct-response telesales because there is nowhere for a bad conversation to hide. A call either ends in an order or it does not.
It also has a hard floor underneath it. Below a certain approval rate the media spend that generated the call does not pay for itself, so work out where that line sits on your own account before the test starts. Hold the result against that number: the agent passes when it clears the line where the account stops losing money.
Give it a calibration period first
Expect the first weeks on live traffic to look like calibration, because that is what they are. The agent handles real callers, gets things wrong, and every failed call goes back into the script and the objection handling the same week.
Judge it once that period is over. What improves an agent over its first weeks is the transcripts: what customers on this particular product actually ask, and the objections nobody wrote down before launch. Scoring those weeks against a floor that has sold the product for years compares a new hire with the veterans.
The calibration weeks are where the script meets the questions nobody wrote down.
Once calibration is over, put the two numbers side by side for the same weeks: the agent's approval rate and the floor's, on the same offers and the same traffic. A different month, a different offer mix or a pilot on leftover calls changes the conditions, and the result then measures those conditions.
Read the result honestly
Whichever side the average favours, do not stop reading there. A floor average blends your best and your weakest people, and the useful question is which of them the agent sells like. Anybody promising that a voice agent will outsell your best people is promising something only your own test can show.
Where it lands among your own operators
A floor average hides the thing that matters, because one floor mixes people who sell very differently. Split your own operators by how well they actually sell, on the same offers in the same weeks, and the comparison starts to mean something:
- Masters, the top of the floor. If the agent matches them, the business case writes itself, so check that result twice before you believe it.
- Mid-level operators, usually the bulk of a floor. Overlapping with this band is a real result, because this is where most of your calls are handled.
- Weak operators. The band an agent has to clear before it takes calls your people would otherwise answer; below it, those calls lose orders a person would have taken.
- Novices, still learning the product. A comparison with them flatters anyone.
- The agent. Place it on this scale offer by offer, because a new offer and a mature one can put it in different bands.
That description is less exciting than most vendor slides and far more useful for planning, because the bulk of a floor sits in the middle band. An agent that sells like that band on every shift, at any hour, is something you can plan around.
The part the conversion number does not show
Conversion is half of the comparison. The other half is how many calls each side can take as they arrive. A floor has a fixed number of seats in any given hour; an agent answers concurrent calls as they come in, up to the capacity agreed for the line. Track calls handled next to approval rate, per offer and per hour, or the test will miss the half where the two differ most.
Media-driven traffic shows why. Calls generated by a TV or radio spot do not arrive evenly. They arrive in a spike when the spot airs, and a floor sized for the spike is idle the rest of the day while a floor sized for the average drops the spike on the ground.
The floor does not lose the spike because it sells badly. It loses the spike because it is finite.
Where a live operator's month actually goes
Before comparing costs, look at what a paid hour on the floor contains. Pull a month of your own workforce reporting and split paid time into:
- Idle time, waiting for a call.
- Time actually in conversation with a customer.
- Breaks and rest.
- Training.
- After-call work.
Conversation is only one entry on that list, and on a floor staffed for a peak nobody can predict, idle time is the entry that grows. That follows from the arithmetic of staffing for a peak, whoever is on the floor. Gartner put the labour share of contact-centre cost at up to 95% in 2022, so an hour spent waiting for the phone is almost all cost and no revenue.
Why we are not printing a cost ratio
A cost ratio between an agent and an operator is built out of local wages, and any ratio we printed would carry somebody else's payroll into your business case. Gartner's own 2025 figures put a self-service contact at 1.84 dollars against 13.50 dollars for a live assisted one, roughly seven to one; use that as the shape and your own payroll for the number.
Where the test applies, and where it does not
A split-traffic test tells you what happens on your line, with your offers and your customers, and nothing more. Short transactional calls are close to the easiest case for automation: a narrow set of questions and a decision made on the call. A considered purchase, a B2B line, or anything where the customer needs a week to decide calls for a test of its own, and a result on one does not carry over to the other.
Nor is the test a headcount plan. A Gartner survey in 2025 found only 20% of service organisations actually cut headcount after deploying AI. Where an agent pays for itself, look first at the calls the floor was dropping at the peak: they show up in revenue.
How to run the test
- Split live traffic, do not run a pilot alongside reality. The only comparison worth anything is the same offers, the same hours, the same customers.
- Give it a calibration period and expect it to be mediocre. Feed every failed call back into the script, and do not score the agent until that period is over.
- Measure the metric the account is actually judged on. Approval rate is unpleasant precisely because it cannot be gamed with a warmer tone.
- Point it at the peak first. The peak is where your floor is losing calls it will never see, and where capacity counts for most.
- Compare against your tiers, not your average. If it sits above your weak band, that is a real result, and your average was never a person anyway.
The question worth answering is narrower than the one on most vendor slides: on your own calls, which of your operators does the agent sell like, and what happens to the calls nobody on the floor was free to take?