Speech recognition, understanding and synthesis run on our own servers, and your call data stays in-region. The agent replies in about 0.3 seconds, so your customer never waits — and when they cut across it, the platform records how much of the reply they actually heard.
People do not notice a good voice. They notice a gap. A pause before the reply is the single quickest way for a caller to work out that nobody is there.
from the caller finishing to the agent starting
Every reply is three steps in sequence: hearing the end of the caller's turn, deciding what to say, and starting to say it. Each one spends milliseconds, and the caller experiences only the total. Running the whole stack ourselves is what makes that total controllable — there is no queue at somebody else's API in the middle of the sentence.
The agent also starts speaking before it has finished composing, the way a person does, so the first word lands while the rest is still being produced.
Stopping when the caller talks is the easy half. The hard half is knowing which part of your sentence they never heard, and doing something about it.
In a recorded stress test of one of our own demo agents, a tester cut across it four times. On the first interruption the platform logged that the caller had heard 76 of the 159 characters the agent was part-way through saying — 47% of the reply. Because the key fact sat in the unheard half, the system cancelled its own move to the next stage, restated the substance, and only then returned to its question.
The other three interruptions landed at 0% heard. The agent acknowledged the caller, repeated the outstanding question and carried on, without talking over them and without losing its place.
The full stress test, including the point we lost →Two failure modes account for most of the calls that go visibly wrong. Both are worth testing before you buy.
A script-driven bot restarts the step. The agent returns to the question it had actually asked, with the answers already collected still in place, so the caller is never taken back to the beginning.
Asked to confirm something that is not in the knowledge base, a model under social pressure will often agree to be helpful. Ours is built to keep two propositions apart rather than collapse them into one convenient answer, and to hand off where a fact is missing.
How we constrain invention →“I am a digital assistant” is honest and, on its own, a dead end. The scenario carries an explicit route out — a callback offer or a transfer — because a caller who asks for a human and is given nothing simply hangs up.
Recognition, understanding and synthesis are ours, in one stack. That is a latency decision and a data decision at the same time.
Speech synthesis, speech recognition, and a language model — the three parts every voice agent needs, running together rather than stitched across vendors.
All three run on our own servers, so your call audio, transcripts and customer data stay in-region and do not pass through third-party services. Where you have a reason to prefer a specific provider for one component, we can plug it in — and we tell you in writing which components are ours and which are not before the agent takes a live call.
Where your data is processed, and what we do not claim →None of this appears in a demo. All of it appears in your call logs.
An answering machine greeting sounds like a person answering, and a pitch delivered to one is pure waste. The agent tells them apart and treats a machine as a machine.
How the detection works →Addresses, spellings, dates and numbers are read back and confirmed on the call, because a delivery to the wrong street costs more than the call did.
Locator reviews 100% of conversations against a checklist you agree with us. A manual QA team reaches 3–5%, which is our own figure from running live call operations.
What call scoring produces →