The demo is not the product
Every vendor demo works. It is supposed to. The script was written by the people who built the agent, the questions are the ones it expects, and the person asking them knows exactly how to phrase them. You are watching a rehearsed performance of the happy path, and every platform on your shortlist has one ready.
That is why demos separate almost nothing. Three vendors will each sound competent for the length of a demo call and you will leave with three impressions and no evidence. The real distance between a phone menu with better speech recognition, a scripted bot, and an agent that can hold a conversation only opens up when the call goes somewhere nobody prepared for.
So stop evaluating the demo. Bring your own script, use the same one with every vendor, and build it out of the calls you actually get: the caller who is confused, the one standing in a car park with the engine running, the one who changes their mind twice, the one who just wants a person. What follows is that script. Ten tests, grouped by when you are able to run them.
Five tests to run while you are still on the call
Start talking while the agent is still speaking, the way people do on the phone every day. Watch whether it stops. A system that finishes reading its paragraph over the top of you has no working interruption handling, and every caller who tries to correct it will end up talking to a wall. This single test eliminates more shortlists than any other.
Say nothing at all for several seconds. Then run the test again from a street, a car, or a room with a television on. You are looking for two failures: an agent that fills the silence by inventing a turn nobody took, and an agent that treats background noise as speech and answers it. Both produce calls that read as unhinged in the transcript.
Point the agent at a number you do not answer and let it run to the recorded greeting. A machine that cannot tell a greeting from a person will hold a full conversation with an answering machine, and you will pay for every minute of it. Ask what happens next, too: whether it leaves a message, whether it schedules a retry, and whether either decision is yours to configure.
Take the call somewhere outside its brief. A dispute, a contract term, a question about a product line it does not cover. There are only two acceptable answers, and both are honest: it says it does not know and hands the call on, or it says it does not know and takes a message. An agent that produces a confident, plausible, invented answer is the one that will do that to a customer next week.
Ask plainly. Then, on a second call, ask badly: mumble it, ask while annoyed, ask sideways with something like "is there anyone else I can speak to". The correct behaviour is to comply on the first clear request, in all the ways people actually phrase it, and to hand over what has already been said so nobody has to repeat themselves. Re-offering self-service to someone who has asked twice is a design decision, and you should know which vendors made it.
That last test is the one buyers routinely under-weight, and the survey evidence says they should not. Most people still want the option of a person on the line, and very few actively prefer the alternative, which means the exit path is not an edge case in your call flow. It is the thing a large share of your callers will reach for the moment the automation stops being useful.
Three tests to run after you hang up
Go and look at the record the call produced. Not the vendor's dashboard, your own system. A call that went beautifully and wrote nothing usable is a call your team cannot act on, and it is a surprisingly common outcome: the conversation is the demo, the data plumbing is the project. Check whether the fields you actually use were filled, and whether the caller's own words survived or were flattened into a category code.
Ask to hear the call back and to read it. Then ask the boring questions: what format, how long it is retained, who can reach it, and whether the caller was told it was being recorded. If a recording cannot be produced on request during a sales evaluation, it will not appear when you need it for a complaint.
Every vendor will say quality is monitored. Ask what that means in practice, and specifically whether it means a sample listened to by a person or every call scored automatically against criteria you set. The difference decides whether you find out about a broken conversation this week or next quarter.
That last question matters more than it sounds, because manual review does not scale and never has. In our own operations, hand-reviewed quality assurance reached under 5% of recorded calls, which is enough to spot a catastrophe and nowhere near enough to spot a pattern. Automated scoring covers 100% of calls against the criteria you define, which is a different kind of instrument: not a spot check, a measurement.
Two tests to run on the contract, not the phone
Ask to see the exact disclosure wording the agent uses, how an opt-out is captured mid-call, how a do-not-call request propagates to the rest of your systems, and where the consent record is stored. Under GDPR, UK GDPR and CCPA the obligation lands on you as the business making the call, not on the vendor supplying the voice, so a vendor who cannot show you the trail is selling you their risk as well as their product.
Demos are priced by the vendor's arithmetic, on the vendor's volume. Redo it on yours. Ask what the unit is, whether it is a minute or a seat or a resolved conversation, what happens at the volumes you hit in a busy week rather than an average one, and what the pilot itself costs before anything is committed. A per-minute price is only comparable to another per-minute price once you know what is inside the minute.
The pricing test is also where the comparison stops being about technology. Two platforms that both pass all nine of the earlier tests can still be a factor apart on the line item that lands in your budget, and the only way to see that is to run your own call volumes through both. If it helps, our own per-minute rates and volume discounts are published rather than quoted on request, and the calculator below will take your numbers instead of ours.
What these 10 tests actually separate
Run this script across a shortlist and the market's three product categories reappear on their own, without anybody's marketing having to be believed. An interactive voice response system routes calls and fails the first two tests immediately, because it was never built to be interrupted or to hear anything but a menu choice. A callbot or a voicebot handles the calls it was scripted for and comes apart on the fourth. What survives the whole script is a different class of thing, and the word a vendor uses for it matters far less than which tests it passed in front of you.
The reframe
You are not buying a voice. You are buying what happens when the call leaves the script, because that is the part your callers will experience and remember. Every test above is a way of paying to find that out now, in an evaluation, instead of later, on a customer.
Write the answers down as you go, one column per vendor, and keep the script identical across all of them. It is not a sophisticated method and it does not need to be. Two or three of the ten tests will usually settle the decision on their own, and you will have evidence for it rather than an impression of a good demo.