Quality Control · 25 August 2026 · 9 min read

AI receptionist versus a clinic front desk: we called two clinics as patients

We called a dental practice and a fertility centre as patients, put the same questions to the front desk and to the neuro-agent on the same account, and scored both out of sixteen. Nobody on either line was rude or badly trained. Both lines still lost the booking, and they lost it on speed, on half-quoted prices and on the fact that nothing gets written down.

OUR OWN MYSTERY-SHOPPER CALLS9/16Fertility clinic front desk16/16Neuro-agent, same accountscored againstOur own scoring of one mystery-shopper call to an anonymised Moscowfertility clinic, eight criteria at zero to two. The front desk knew the medicineand lost on speed, transfers and the fact that nothing was written down. Thedental clinic on the same test finished 4/16 against 15/16.

A patient who wants an implant does not call one clinic. They call four, in the same hour, from the same sofa, and the shortlist gets made out of whichever calls went well. That is the part most clinics never see: the comparison is happening on the phone, and the phone is usually the least managed thing they own. So we called two of our own clients as patients and scored what happened.

How we measured

Two anonymised private clinics in Moscow, both of them accounts of ours: one dental practice, one fertility centre. We put the same request to the clinic's own front desk and to the neuro-agent running on the same account, and scored both against eight criteria at zero to two. These are our own recordings and our own scores, not a study. We are not publishing the recordings and we are not naming the clinics. Nothing below is a price: what crosses over from a Moscow price list to your market is the shape of the answer, not the number.

Call one: what does a turnkey implant cost

The question was the one every implant patient asks first and the one almost every clinic answers worst: what does the whole thing cost, start to finish, with nothing left to add later.

Getting to a human took 53 seconds. Then the administrator put the call on hold for another 34 seconds to fetch somebody who could talk about prices. One minute and 41 seconds before the question could even be asked. The neuro-agent on the same account picked up in 2 seconds.

Then the interesting part. Asked what a turnkey implant costs, the front desk named a figure that came to about half the clinic's own turnkey total. Asked again, in the same call, it gave the same half.

Asked twice for the whole price, the front desk quoted about half of it twice.

That is worse than quoting high. A patient who hears half the number puts the clinic at the top of the shortlist, arrives, finds out the abutment and the crown are separate lines, and leaves feeling handled. The clinic wins the call and loses the patient, which is the most expensive way there is to lose one.

The agent answered differently on the same price list, on the same day. It named the four things a turnkey implant is actually made of, said plainly that the final figure depends on an examination, and offered to book the examination.

What the caller is really doing

They are not asking for a number. They are asking for a number they can compare. An incomplete answer is not a cheaper answer, it is an unusable one, and the caller finds that out at the next clinic on the list rather than at yours.

Call two: can you even reach a coordinator

The fertility centre had the opposite problem. The people were better and the line was worse.

Across the calls, the average wait before a human voice said anything was about 100 seconds. Hold music ran from 1:19 to 2:23 depending on the call. One of the calls ended without the coordinator ever being reached at all.

When a call did land, it took 10 minutes and 2 transfers, and the caller explained the same situation three times to three different people. The agent on the same account took 4 minutes, in one unbroken conversation, and asked nothing twice.

Instead of an answer on programme eligibility, the live line offered a callback inside 15 minutes. The callback was not the problem. The problem showed up the next day: nothing from the first call had been written down, so the second call started from nothing. On that follow-up the live line scored 1/6.

The cheapest thing a clinic can do on the phone is remember the last call. Almost nobody does it.

The scorecard

Both calls were scored the same way, on eight things a caller notices whether or not they could name them:

Each one scored zero, one or two. The dental front desk finished on 4/16 against the agent's 15/16. The fertility front desk, which was genuinely good at the medicine, finished on 9/16 against 16/16.

The uncomfortable part

Neither front desk was rude, lazy or badly trained. The fertility coordinator handled every clinical question and every objection correctly and still lost most of the available points, because most of the points are operational rather than clinical.

What both calls have in common

Three failures, and not one of them is about attitude:

  1. Speed. Not the length of the call, the time before the caller hears a human being. ContactBabel put the average across UK contact centres at 116 seconds in 2024, and a Velaro survey back in 2012 found around 60% of callers will not hold longer than a minute. Both of these clinics sit inside that trap.
  2. Completeness. The front desk knows the price list exists. It does not have the price list in front of it, so it quotes the part it remembers, which is the first line and the smallest one.
  3. Memory. Nothing said during a call survives the call. Every later conversation restarts, and the caller reads that as being a stranger to a business they have already spoken to twice.

None of the three gets fixed by a training session, because none of the three is a knowledge problem. They are all what happens to a person under load with no reliable place to put things.

What the AI voice agent actually changed

On both accounts the neuro-agent answered in 2 seconds, quoted the structure of the price rather than its first line, captured the caller's details, and put what it had learned into the record before the call ended.

It is worth being precise about why, because it is not intelligence. The agent has the price list open on every single call, and it cannot decide that this particular caller probably only wants the headline figure. It does not get tired at the end of a shift. It does not forget to write the note, because writing the note is not a separate task it has to remember to do.

The agent did not out-think anyone on that line. It simply never forgot the price list.

What it did not do

It did not diagnose anything and it did not try. On the dental call it scored 15 out of 16 rather than 16, and the missing point is the honest one: it could not give a final price either, because a final price needs an examination. What it could do was say exactly that, in one sentence, instead of quoting half a number twice.

It also did not replace the coordinator. On the fertility account the clinical conversation still belongs to the coordinator, and it is better than anything an agent will do this year. What changed is that the coordinator now picks up a call with the intake already taken. That is the arrangement we argue for everywhere: the agent takes the part that is arithmetic and the human takes the part that is judgement.

What to do with your own line

You can run this test on yourself this week, and you probably should, because the result is usually worse than anyone in the building expects.

  1. Call your own main number from a phone nobody there recognises, at a normal busy hour, and time the wait before a human voice says anything.
  2. Ask for the total price of your most expensive service, then ask again in the same call. Compare both answers against your own price list.
  3. Call back the next day, mention the first call, and see whether anybody knows what you are talking about.
  4. Score all of it out of sixteen on the eight criteria above, so the result is a number your team can argue with instead of an opinion they can dismiss.
  5. Fix in this order: answer time, then completeness, then memory. Answer time is the only one your callers are actively timing.

Everything in this piece came out of doing exactly that on two clinics that were confident their phones were fine.

Frequently asked questions about testing your own clinic line

How do you score a phone call out of sixteen?
Eight criteria, each scored zero, one or two: speed of answer, establishing who is on the line, identifying the need, completeness of the answer, whether the price was comparable, handling the objection, ending on a next step, and capturing contact details. Zero means it did not happen, one means partly, two means fully. The point of the scale is not precision, it is that two people listening to the same recording land on roughly the same number.
Are these real calls?
Yes. They are our own recordings on two anonymised client accounts in Moscow, and the scores are ours. We are not publishing the recordings, naming the clinics or reproducing what anyone said word for word. What we publish is our scoring and the timings.
Why not simply train the staff to quote the full price?
Because it is not a knowledge problem. The administrator knew a fuller price existed; the price list was not in front of them and the caller was waiting. Training moves the problem for a few weeks. Having the price list open on every call moves it permanently, which is the actual difference the agent made.
Would an AI voice agent have booked the appointment?
On the dental call it captured the caller's details and offered the examination, which is the next step that call needed. Whether it can put a slot in the diary itself depends on whether the clinic's calendar is connected to it; that is a plumbing question rather than a conversation one.
Do patients mind talking to an AI on a medical line?
Some do, and a Five9 consumer survey in 2024 found 75% of consumers still prefer talking to a real human. That is an argument for saying openly that the caller has reached an agent and for handing over the moment the conversation stops being administrative, not an argument for making them wait about 100 seconds for the human instead.
What is the first thing to fix on a clinic line?
Answer time, every time. A caller who never reaches anybody cannot be impressed by how good your coordinator is. Once somebody answers immediately, the next thing to fix is whether the answer they get is complete enough to compare against the clinic they call next.