AI Agent Technology · 4 August 2026 · 12 min read

Build vs buy an AI voice agent: the technology is the easy part

Build vs buy for an AI voice agent looks like a technology decision and mostly isn't. The speech stack has commoditised. The price matrix, the objection library, the qualification logic and the escalation rules have not, and those are what decide whether the agent earns anything.

BUILD VS BUY VOICE AIThe easy partTechnology: models, latency, hostingThe hard partKnowing what the agent should say to yourcustomersof a voice agent that earnsOur own framing rather than a study: the platform layer is on theshelf for anyone, and the commercial knowledge behind the script iswhat decides whether the agent earns.

The obvious question comes first: why pay per minute when an in-house team could build this? Fair question, and many teams probably could. A competent engineer can wire speech recognition, a language model and a text-to-speech engine into something that answers a phone number, and it will not take them long. What happens after the demo decides whether it earns money.

Here is our framing, labelled as ours rather than dressed up as research: a working voice agent is roughly 30% technology and 70% knowledge of how the business actually sells. The 30% is the stack, the latency, the hosting. The 70% is the price matrix, the objection library, the qualification logic and the escalation rules, which decide what the agent says when a customer pushes back on price. Buy the model without the knowledge and you own a capable car with nobody in the driver's seat.

None of that means you should never build. Sometimes building is right, and we will say exactly when. But most teams answer the build vs buy question by comparing the wrong two things: a one-off engineering project against a recurring invoice.

What actually commoditised

A few years ago, putting a passable synthetic voice on a phone line was a research problem. It isn't now. Open-weight language models, hosted recognition and synthesis, telephony APIs that take a few lines of setup: the pieces are on the shelf and cheap to try. A competent backend developer can produce a convincing demo quickly.

We will say the uncomfortable part out loud, since hiding it works against our own interest. The platform layer is not a moat. We run recognition, synthesis and the language model on our own servers, which buys replies in about 0.3 seconds and keeps call data in-region without depending on third-party APIs. But a serious team with time and budget can reproduce that.

A demo that answers the phone is a small piece of work. An agent that closes a sale is somebody's domain knowledge, written down precisely enough for software to follow.

What the technology alone buys you

A voice that can hold a coherent conversation, and no idea what it should say to a customer who thinks the price is too high.

The knowledge nobody sells you

The part we are calling the 70%, again our own framing rather than a study, is not soft skills or culture. It is a set of concrete artefacts, and without them in writing the agent has nothing to run on.

Check whether any of this exists on paper in your own company. Where it does not, it exists only in what your best salespeople know and in recordings nobody has listened to. That gap, rather than the model, is the one we would expect a build to stall on.

It is also why we tell clients to start with speech analytics. Locator scores 100% of calls; a manual quality team gets through 3 to 5%, our own figure rather than an industry study. The objection library comes out of those transcripts, not out of anyone's imagination.

When building in-house is the right call

There is a real build case, and it deserves stating fairly. Building is right when several of these hold at once.

  1. Your engineering team already owns the telephony. If your developers run the dialler, the routing and the recording pipeline, the hardest infrastructure is behind you.
  2. The use case is narrow and stable. An agent that confirms a delivery slot or reads back a balance has a script that survives every promotion untouched.
  3. Volume is large enough that per-minute pricing stops being the cheap option. This one is arithmetic, not opinion.
  4. The commercial knowledge is already written down. A company with a maintained price matrix and a real objection library is buying only the technology layer, which is a much better reason to build.

On volume, here is the arithmetic on our published rates. Full-cycle selling lists at EUR 0.30 a minute excluding VAT, the same nominal figure in USD for the US. Volume tiers cut it: 5% from 50,000 minutes a month, 10% from 100,000, 15% from 200,000, 20% from 400,000, 25% from 1 million. Billing covers talk time only, analytics included, no subscription.

25,000 minutes at EUR 0.300%
Base list rate, no volume discount, excluding VAT.
100,000 minutes at EUR 0.270%
Published 10% volume tier, excluding VAT.
400,000 minutes at EUR 0.240%
Published 20% volume tier, excluding VAT.
1,000,000 minutes at EUR 0.2250%
Published 25% volume tier, excluding VAT — about EUR 2.7 million a year.

Set those against your own build estimate, and make the estimate include the people, not their first sprint. EUR 2.7 million a year buys a lot of engineering, and a company at that volume with stable requirements should model a build seriously. At the bottom of the table the whole bill is EUR 7,500 a month, and the build case there has to survive a comparison with the loaded monthly cost of the people who would own it — a figure only you can fill in.

The honest build case is not "we could make this". It is "we will still be maintaining it in two years, deliberately, with named people".

The maintenance bill nobody budgets for

A build gets budgeted as a project with an end date. A voice agent is not a project. It degrades quietly whenever the business around it moves.

None of this is exotic. It is ordinary product maintenance. The point is that it lands on a headcount line, not a project line.

What to demand from a vendor before you sign

If you buy, the protection isn't a longer contract. It is a short list of questions asked before signing.

  1. Real recordings from your category, including calls that went badly. A scripted demo tells you about the demo.
  2. A written per-minute price at your expected volume, with the tier and discount stated. A price that needs a meeting to disclose needs a meeting to change.
  3. An exact list of what is billable. Does ringing, hold time or answering-machine time reach the invoice? We bill talk time only and include analytics.
  4. Where the models run and where audio goes. Behind someone else's API, your call data and your response time depend on a third party.
  5. A trial you can leave. Ours is priced 20% above the base rate, and with no subscription there is nothing to unwind, which is the honest shape of optionality.
  6. A named owner for the commercial knowledge. Who writes the objection library, who updates the price matrix after go-live, and how fast does a change land? If that owner is you, you bought the technology only.

How we split the work with clients

Our own answer isn't "buy everything from us". The two halves belong to different owners, and saying so early prevents most of the disappointment.

01
You bring the commercial truth

Price bands and discount authority, what counts as a qualified lead, who gets escalated and when. Nobody outside your company can invent this.

02
We read your calls first

Locator scores 100% of your calls at EUR 0.10 a minute excluding VAT, so the objection library comes out of your own calls rather than anyone's imagination. That is also how you find out whether your commercial knowledge is consistent enough to write down.

03
We build, host and run the agent

Recognition, synthesis and the language model on our servers, replies in about 0.3 seconds, data staying in-region. The technology layer, and not worth rediscovering in-house without a reason.

04
We move down the risk ladder in order

Speech analytics, then the inbound line agent, then the qualifier, then the full-cycle seller. Each stage stands alone, so you can stop after any of them without stranding the work.

The direction we are moving in is pay-for-results: per sale, per qualified lead, per completed action. It only works if the vendor carries the commercial knowledge too, since nobody gets paid for an excellent agent that fails to sell. An in-house team carries that exposure with no counterparty.

Build vs buy an AI voice agent: the short version

Build if your engineers already own the telephony, the use case is narrow and unlikely to change, your volume turns the per-minute arithmetic against buying, and your commercial knowledge is documented. If two or more of those are false, buying is the lower-risk option, and the gap you would have to close first is not the technology.

Either way, do the arithmetic before the architecture. Our rates and volume tiers are on the homepage pricing section, and the cost calculator turns your call volume, your region and the product you want into a monthly figure you can set next to a build estimate.

Frequently asked questions

Should we build our own AI voice agent or buy one?
Buy unless several things are true at once: your engineering team already owns the telephony, the use case is narrow and stable, your monthly volume is high enough that per-minute pricing turns against you, and your price matrix and objection library are already written down. If two or more of those are missing, the harder gap to close is the missing commercial knowledge, not the technology.
Why do you say a voice agent is only 30% technology?
The split is our own framing rather than a measured study: we put the stack, the latency and the hosting at roughly 30% of what makes an agent work, and that part is now available to anyone. The other 70% is knowing what the agent should say — the price matrix, the objection library, the qualification logic and the escalation rules. Buy the model without that knowledge and you own a capable car with nobody in the driver's seat.
What is in the 70% you say is not technology?
On our own framing rather than any study, the non-technology part is five artefacts, all of which have to exist in writing. The price matrix, meaning which discount may be offered, to whom, when, and in exchange for what. The objection library in your customers' own words. The qualification logic, precise enough for software to apply mid-call. The escalation rules. And a written definition of a good outcome, so calls can be scored against something.
At what call volume does building in-house start to make sense?
Check the arithmetic on our published rates. Full-cycle selling is EUR 0.30 a minute excluding VAT, so 25,000 minutes a month is EUR 7,500. At 1 million minutes the 25% volume discount brings the rate to EUR 0.225, which is EUR 225,000 a month, about EUR 2.7 million a year. That second figure buys real engineering capacity, and at that scale with stable requirements a build deserves serious modelling. The first has to be weighed against the loaded monthly cost of the people who would own the build, which is a figure only you can supply. Latin America is priced at roughly half the Europe and US rates, so the threshold sits higher there.
What should we ask a voice AI vendor before signing?
Ask for real recordings in your category, including calls that went badly. Ask for a written per-minute price at your expected volume with the discount tier stated. Ask exactly what is billable, since ringing, hold time and answering machines can all reach an invoice. Ask where the models run and where audio is stored. Ask for a trial you can exit. And ask who owns the objection library and the price matrix after go-live, by name.
What's the cheapest way to find out whether that commercial knowledge exists?
Have speech analytics score your calls. Locator is EUR 0.10 a minute excluding VAT and scores 100% of calls, against the 3 to 5% a manual quality team gets through, which is our own claim rather than a published industry study. Once enough of your calls have been scored, you will be able to see whether your objections, your working replies and your qualification rules are consistent enough to write down, which is the same thing as knowing whether an agent could run on them.