The obvious question comes first: why pay per minute when an in-house team could build this? Fair question, and many teams probably could. A competent engineer can wire speech recognition, a language model and a text-to-speech engine into something that answers a phone number, and it will not take them long. What happens after the demo decides whether it earns money.
Here is our framing, labelled as ours rather than dressed up as research: a working voice agent is roughly 30% technology and 70% knowledge of how the business actually sells. The 30% is the stack, the latency, the hosting. The 70% is the price matrix, the objection library, the qualification logic and the escalation rules, which decide what the agent says when a customer pushes back on price. Buy the model without the knowledge and you own a capable car with nobody in the driver's seat.
None of that means you should never build. Sometimes building is right, and we will say exactly when. But most teams answer the build vs buy question by comparing the wrong two things: a one-off engineering project against a recurring invoice.
What actually commoditised
A few years ago, putting a passable synthetic voice on a phone line was a research problem. It isn't now. Open-weight language models, hosted recognition and synthesis, telephony APIs that take a few lines of setup: the pieces are on the shelf and cheap to try. A competent backend developer can produce a convincing demo quickly.
We will say the uncomfortable part out loud, since hiding it works against our own interest. The platform layer is not a moat. We run recognition, synthesis and the language model on our own servers, which buys replies in about 0.3 seconds and keeps call data in-region without depending on third-party APIs. But a serious team with time and budget can reproduce that.
A demo that answers the phone is a small piece of work. An agent that closes a sale is somebody's domain knowledge, written down precisely enough for software to follow.
What the technology alone buys you
A voice that can hold a coherent conversation, and no idea what it should say to a customer who thinks the price is too high.
The knowledge nobody sells you
The part we are calling the 70%, again our own framing rather than a study, is not soft skills or culture. It is a set of concrete artefacts, and without them in writing the agent has nothing to run on.
- The price matrix. Not a price list. Which discount the agent may offer, to whom, at what point, and for what in return.
- The objection library. The sentences your customers actually use, with replies that work on your product. Generic rebuttals are useless here.
- The qualification logic. What makes a lead worth a person's time, defined precisely enough for software to decide it mid-call.
- The escalation rules. When the agent must stop and hand the call to a human, and what it says while doing it.
- A definition of a good outcome. A sale, a booked appointment, a confirmed address, an agreed callback. Without one, nothing can be scored.
Check whether any of this exists on paper in your own company. Where it does not, it exists only in what your best salespeople know and in recordings nobody has listened to. That gap, rather than the model, is the one we would expect a build to stall on.
It is also why we tell clients to start with speech analytics. Locator scores 100% of calls; a manual quality team gets through 3 to 5%, our own figure rather than an industry study. The objection library comes out of those transcripts, not out of anyone's imagination.
When building in-house is the right call
There is a real build case, and it deserves stating fairly. Building is right when several of these hold at once.
- Your engineering team already owns the telephony. If your developers run the dialler, the routing and the recording pipeline, the hardest infrastructure is behind you.
- The use case is narrow and stable. An agent that confirms a delivery slot or reads back a balance has a script that survives every promotion untouched.
- Volume is large enough that per-minute pricing stops being the cheap option. This one is arithmetic, not opinion.
- The commercial knowledge is already written down. A company with a maintained price matrix and a real objection library is buying only the technology layer, which is a much better reason to build.
On volume, here is the arithmetic on our published rates. Full-cycle selling lists at EUR 0.30 a minute excluding VAT, the same nominal figure in USD for the US. Volume tiers cut it: 5% from 50,000 minutes a month, 10% from 100,000, 15% from 200,000, 20% from 400,000, 25% from 1 million. Billing covers talk time only, analytics included, no subscription.
Set those against your own build estimate, and make the estimate include the people, not their first sprint. EUR 2.7 million a year buys a lot of engineering, and a company at that volume with stable requirements should model a build seriously. At the bottom of the table the whole bill is EUR 7,500 a month, and the build case there has to survive a comparison with the loaded monthly cost of the people who would own it — a figure only you can fill in.
The honest build case is not "we could make this". It is "we will still be maintaining it in two years, deliberately, with named people".
The maintenance bill nobody budgets for
A build gets budgeted as a project with an end date. A voice agent is not a project. It degrades quietly whenever the business around it moves.
- Commercial changes. Every price change, promotion and revised delivery term is a script change, due the same week.
- New objections. Competitors launch and customers find new reasons to say no. Catching that means somebody scores calls every week.
- Latency. About 0.3 seconds to a reply is not unlocked once. You defend it against every model change and every busy hour.
- Model upgrades. A newer model phrases things differently, sometimes better, sometimes in a way that breaks the one step that used to convert.
- The dull mechanics. Spotting an answering machine, handling interruptions mid-sentence, confirming an address without three repeats. Unglamorous, and obvious to the customer when missing.
- Data and legal. Where audio is processed and stored is a standing responsibility. We are registered in Dubai and use an EU representative for EU data.
None of this is exotic. It is ordinary product maintenance. The point is that it lands on a headcount line, not a project line.
What to demand from a vendor before you sign
If you buy, the protection isn't a longer contract. It is a short list of questions asked before signing.
- Real recordings from your category, including calls that went badly. A scripted demo tells you about the demo.
- A written per-minute price at your expected volume, with the tier and discount stated. A price that needs a meeting to disclose needs a meeting to change.
- An exact list of what is billable. Does ringing, hold time or answering-machine time reach the invoice? We bill talk time only and include analytics.
- Where the models run and where audio goes. Behind someone else's API, your call data and your response time depend on a third party.
- A trial you can leave. Ours is priced 20% above the base rate, and with no subscription there is nothing to unwind, which is the honest shape of optionality.
- A named owner for the commercial knowledge. Who writes the objection library, who updates the price matrix after go-live, and how fast does a change land? If that owner is you, you bought the technology only.
How we split the work with clients
Our own answer isn't "buy everything from us". The two halves belong to different owners, and saying so early prevents most of the disappointment.
Price bands and discount authority, what counts as a qualified lead, who gets escalated and when. Nobody outside your company can invent this.
Locator scores 100% of your calls at EUR 0.10 a minute excluding VAT, so the objection library comes out of your own calls rather than anyone's imagination. That is also how you find out whether your commercial knowledge is consistent enough to write down.
Recognition, synthesis and the language model on our servers, replies in about 0.3 seconds, data staying in-region. The technology layer, and not worth rediscovering in-house without a reason.
Speech analytics, then the inbound line agent, then the qualifier, then the full-cycle seller. Each stage stands alone, so you can stop after any of them without stranding the work.
The direction we are moving in is pay-for-results: per sale, per qualified lead, per completed action. It only works if the vendor carries the commercial knowledge too, since nobody gets paid for an excellent agent that fails to sell. An in-house team carries that exposure with no counterparty.
Build vs buy an AI voice agent: the short version
Build if your engineers already own the telephony, the use case is narrow and unlikely to change, your volume turns the per-minute arithmetic against buying, and your commercial knowledge is documented. If two or more of those are false, buying is the lower-risk option, and the gap you would have to close first is not the technology.
Either way, do the arithmetic before the architecture. Our rates and volume tiers are on the homepage pricing section, and the cost calculator turns your call volume, your region and the product you want into a monthly figure you can set next to a build estimate.