Quality Control · 6 September 2026 · 7 min read

How to stress-test an AI voice agent before you buy one

Every vendor demo is staged. Here is the five-criterion test you can run yourself in a single call, the three probes that break an agent fastest, and the scorecard our own agent came back with when we ran it on ourselves.

FIVE-CRITERION TEST1Response latency2Barge-in resilience3Context retention4Hallucination probe5Human handoffFour criteria held. The fifth, giving the caller a route to a person, is whereours lost its point.

A vendor demo is chosen ground. The scenario is picked in advance, the questions are known, and nobody in the room is trying to break anything. You come away knowing what the agent does on its best day, which is the one thing you did not need to find out.

What you actually need is the opposite exercise: a short, deliberately awkward call, scored against criteria you set before you dial. It takes one phone call and a notepad. You do not need the vendor's permission and you do not need a trial account.

The five things worth scoring

A sales-quality checklist tells you whether the agent sold well. It tells you nothing about whether the agent holds together when the call goes sideways, and those are separate tests. This is the technical one: five criteria, two points each.

A demo tells you what the agent does when everything goes right. A stress test tells you what it does when the caller is difficult, which is most callers.

Probe one: cut across it four times

Interrupt the agent at four different points: mid-sentence on its opening, halfway through an answer, immediately after it asks you something, and once while it is reading back a detail. Handling this well is harder than it sounds, because the agent has to answer three questions at once. Has the caller started speaking? How much of my reply did they actually hear? And what do I do about the part they missed?

That middle question is the one most systems skip entirely, and it is the one to listen for. When we ran this probe against a demo agent of our own, the platform recorded that on the first interruption the caller had heard 76 of the 159 characters the agent was part-way through saying — 47% of the reply. The key detail sat in the unheard half, so the agent cancelled its own move to the next stage, restated the substance, and only then returned to its question.

The three later interruptions came earlier in the turn: 0% heard each time. The agent acknowledged the caller, repeated the outstanding question, and carried on without ever talking over them or losing its place in the sequence. That is what a pass looks like. A fail sounds like an agent that stops politely and then continues as though the unheard half had landed.

The metric to ask your vendor for

Not “does it handle interruptions” — every vendor says yes. Ask whether the platform records how much of each reply the caller actually heard before cutting in, and whether the agent uses that number to decide what to repeat. That is the difference between stopping politely and understanding what was missed.

Probe two: ask about someone who does not exist

Invent a person, a product tier or a service the company does not offer, and ask about it as though someone else had already told you it existed. Social pressure plus a false premise is a reliable way to get a model to agree with something untrue, because agreeing is the helpful-sounding move.

There are two failure modes and you are listening for both. The agent that flatly denies the thing exists when it partly does, and the agent that confirms whatever you handed it. A pass keeps the two propositions apart: it says what is actually true, says plainly that the rest does not match anything it has, and offers a real alternative instead of a comfortable answer.

Probe three: ask for a person

Ask directly whether you are talking to a machine and whether you can speak to a human being instead. This is the cheapest probe to run and the one most agents are quietly worst at.

There are three distinct ways to fail it: claiming to be human, pretending to transfer a call it cannot transfer, and simply ignoring the question. Avoiding all three is not automatic. But avoiding them is only half the criterion, and this is the point our own agent lost. It answered honestly, described what it could help with, did not fake a transfer — and then offered nothing else. No mention that a person handles the appointment itself, no offer of a callback, no route out.

That scored one point of the two, and it is the finding worth publishing about ourselves. Honesty about being a machine is necessary and not sufficient. For some callers “I am a digital assistant” closes the question; for others it reads as a dead end, and a dead end at the moment someone is asking for help is expensive in a way no dashboard shows you. The caller does not complain. They hang up and phone a competitor, and the transcript records a polite, correct, unsuccessful conversation.

Telling the caller you are a machine is the easy half. Giving them somewhere to go next is the half that keeps the call.

What our own agent scored

We run this test against our own agents rather than only against the ones we are comparing ourselves to. Here is the most recent scorecard in full, including the criterion we failed, because a vendor who will only show you the four passes is telling you something about the fifth.

That is 9 out of 10, and the nine is not the interesting part. The missing point is, because it names a fix: an explicit exit path written into the scenario, either a sentence confirming that a person handles the next step or a straightforward offer of a callback. That is a scripting decision, not a limit of the model, and it only surfaced because somebody deliberately went looking for it.

Running the test yourself

You do not need our checklist to do this. You need one call, a willingness to be awkward for four minutes, and a note of what happened.

  1. Write the five criteria down before you dial, and decide what a pass sounds like for each one. Scoring after the fact is how you talk yourself into a result.
  2. Call the vendor's own demo line, not a scripted walkthrough with a salesperson listening.
  3. Run the three probes in order: interruptions first, then the false premise, then the request for a person.
  4. Score each criterion out of two the moment you hang up, while you can still hear it.
  5. Ask the vendor for the scorecard from their own last stress test, and for the criterion they failed.

That last step is the one that separates vendors. A supplier who cannot produce a failed criterion has either never run the test or would rather you did not see the result. Either answer tells you what you are buying.

What happens after you buy

A stress test is a snapshot of one call. Once an agent is live the question becomes how many of the real ones get checked, and by whom — which is a different job from choosing a vendor, and one worth asking about in the same conversation.

Frequently asked questions

How do you stress-test an AI voice agent?
Run one deliberately awkward call against five criteria you write down before you dial: response latency, barge-in resilience, context retention, hallucination provocation and human handoff, two points each. Three probes cover all five — interrupt the agent four times, ask about something that does not exist, and ask to speak to a person.
What should an AI voice agent do when a caller interrupts it?
Stop cleanly, work out how much of its reply the caller actually heard, and repeat the part that was missed before moving on. The measurement is what matters: on one of our own logged calls the caller had heard 76 of the 159 characters the agent was saying, so the agent restated the substance instead of advancing the script.
What should you ask a voice agent vendor before buying?
Ask for the scorecard from their own most recent stress test and for the criterion they failed. Ask whether the platform logs how much of each reply a caller heard before interrupting. A vendor who cannot produce a failed criterion has either never run the test or would rather you did not see it.