Call Center Economics · 25 August 2026 · 9 min read

AI voice agent versus a live sales floor: how to run a fair test

A demo shows that an AI voice agent can talk; only live traffic shows whether it sells like your operators. Here is how to split real calls between the agent and your sales floor, what to compare it against, and where to look for the money besides the conversion rate.

A FAIR TESTA demoShows the agent can hold a conversationLive trafficShows whether it sells like your operatorsversusOur own framing for a fair test: split real calls between the floor andthe agent on the same offers, score both on approval rate, and placethe agent among your operator tiers.

Most claims about voice agents are made against a demo: a scripted caller, a happy path, a recording picked afterwards. A demo shows that a neuro-agent can hold a conversation. The owner of a sales line needs a second answer: whether it sells as well as the people already on the phones.

Only a comparison on live traffic gives that answer. It is harder to set up than a demo and can embarrass either side, which is exactly why it is worth running.

A method for your own test

This article sets out how to run the comparison and how to read it. It reports no conversion rate, headcount or cost from any account of ours, because a number from one line, one product and one market would tell you little about yours. The only figures below are Gartner's, each with its year.

What approval rate is, and why it is a brutal metric

Approval rate is confirmed orders over calls handled. Not leads, not interest, not a callback agreed: an order the customer actually confirmed on that call. It is the least forgiving number in direct-response telesales because there is nowhere for a bad conversation to hide. A call either ends in an order or it does not.

It also has a hard floor underneath it. Below a certain approval rate the media spend that generated the call does not pay for itself, so work out where that line sits on your own account before the test starts. Hold the result against that number: the agent passes when it clears the line where the account stops losing money.

Give it a calibration period first

Expect the first weeks on live traffic to look like calibration, because that is what they are. The agent handles real callers, gets things wrong, and every failed call goes back into the script and the objection handling the same week.

Judge it once that period is over. What improves an agent over its first weeks is the transcripts: what customers on this particular product actually ask, and the objections nobody wrote down before launch. Scoring those weeks against a floor that has sold the product for years compares a new hire with the veterans.

The calibration weeks are where the script meets the questions nobody wrote down.

Once calibration is over, put the two numbers side by side for the same weeks: the agent's approval rate and the floor's, on the same offers and the same traffic. A different month, a different offer mix or a pilot on leftover calls changes the conditions, and the result then measures those conditions.

Read the result honestly

Whichever side the average favours, do not stop reading there. A floor average blends your best and your weakest people, and the useful question is which of them the agent sells like. Anybody promising that a voice agent will outsell your best people is promising something only your own test can show.

Where it lands among your own operators

A floor average hides the thing that matters, because one floor mixes people who sell very differently. Split your own operators by how well they actually sell, on the same offers in the same weeks, and the comparison starts to mean something:

That description is less exciting than most vendor slides and far more useful for planning, because the bulk of a floor sits in the middle band. An agent that sells like that band on every shift, at any hour, is something you can plan around.

The part the conversion number does not show

Conversion is half of the comparison. The other half is how many calls each side can take as they arrive. A floor has a fixed number of seats in any given hour; an agent answers concurrent calls as they come in, up to the capacity agreed for the line. Track calls handled next to approval rate, per offer and per hour, or the test will miss the half where the two differ most.

Media-driven traffic shows why. Calls generated by a TV or radio spot do not arrive evenly. They arrive in a spike when the spot airs, and a floor sized for the spike is idle the rest of the day while a floor sized for the average drops the spike on the ground.

The floor does not lose the spike because it sells badly. It loses the spike because it is finite.

Where a live operator's month actually goes

Before comparing costs, look at what a paid hour on the floor contains. Pull a month of your own workforce reporting and split paid time into:

Conversation is only one entry on that list, and on a floor staffed for a peak nobody can predict, idle time is the entry that grows. That follows from the arithmetic of staffing for a peak, whoever is on the floor. Gartner put the labour share of contact-centre cost at up to 95% in 2022, so an hour spent waiting for the phone is almost all cost and no revenue.

Why we are not printing a cost ratio

A cost ratio between an agent and an operator is built out of local wages, and any ratio we printed would carry somebody else's payroll into your business case. Gartner's own 2025 figures put a self-service contact at 1.84 dollars against 13.50 dollars for a live assisted one, roughly seven to one; use that as the shape and your own payroll for the number.

Where the test applies, and where it does not

A split-traffic test tells you what happens on your line, with your offers and your customers, and nothing more. Short transactional calls are close to the easiest case for automation: a narrow set of questions and a decision made on the call. A considered purchase, a B2B line, or anything where the customer needs a week to decide calls for a test of its own, and a result on one does not carry over to the other.

Nor is the test a headcount plan. A Gartner survey in 2025 found only 20% of service organisations actually cut headcount after deploying AI. Where an agent pays for itself, look first at the calls the floor was dropping at the peak: they show up in revenue.

How to run the test

  1. Split live traffic, do not run a pilot alongside reality. The only comparison worth anything is the same offers, the same hours, the same customers.
  2. Give it a calibration period and expect it to be mediocre. Feed every failed call back into the script, and do not score the agent until that period is over.
  3. Measure the metric the account is actually judged on. Approval rate is unpleasant precisely because it cannot be gamed with a warmer tone.
  4. Point it at the peak first. The peak is where your floor is losing calls it will never see, and where capacity counts for most.
  5. Compare against your tiers, not your average. If it sits above your weak band, that is a real result, and your average was never a person anyway.

The question worth answering is narrower than the one on most vendor slides: on your own calls, which of your operators does the agent sell like, and what happens to the calls nobody on the floor was free to take?

Frequently asked questions about voice agents and conversion

Does an AI voice agent convert better than human operators?
Nobody can answer that for your line without testing it there. Split live calls between the agent and your floor on the same offers, then compare the agent with each of your operator tiers: clearing your weak and novice bands is a real result, and matching your best people is a claim to check twice.
What is approval rate?
Confirmed orders divided by calls handled, on a direct-response inbound line, where an order is one the customer agreed to on the call. It is the metric this kind of account lives on, and below a threshold each account has to work out for itself, the media spend that generated the call stops paying for itself.
Why give the agent a calibration period?
Because the first weeks on live traffic are when the script meets the questions nobody wrote down. Every failed call goes back into the script and the objection handling; scoring the agent before that is done compares a new hire with operators who have sold the product for years.
Why is there no cost saving figure in this article?
Because a cost ratio between an agent and an operator is built out of local wages and would mislead anyone whose payroll differs. Gartner's 2025 figures put a self-service contact at 1.84 dollars against 13.50 dollars for a live assisted one, about seven to one; that is the shape of the saving. Your own payroll sets its size.
Does an AI voice agent mean we can reduce headcount?
Do not plan on it: a Gartner survey in 2025 found only 20% of service organisations actually cut headcount after deploying AI. Look for the gain in the traffic peak that currently goes unanswered, which shows up as revenue. Plan around the calls you are currently losing.
Would this work on a considered purchase or in B2B?
Not without a test of its own. Short transactional calls are close to the easiest case there is: a narrow set of questions and a decision made during the call. A long sales cycle with several stakeholders is a different problem; there an agent is most useful on qualification and follow-up, and closing stays with your people.