Agents that hold a phone conversation, at volume, without sounding like a phone tree.
Voice is harder than chat. Interruptions, accents, silence, someone talking over the agent — it has to handle all of it in under a second, or the caller hangs up.
We have built and run a voice agent in production at interview volume. That product is Vaani, and it is why we can talk about latency and barge-in from experience rather than a vendor datasheet.
- Response has to start before the caller thinks it has failed.
- The caller will interrupt. The agent has to stop and listen.
- Accents, background noise and bad lines are the normal case.
- A wrong answer said confidently is worse on a call than in text.
- Recordings are personal data from the first second.
Calls worth automating
Structured first-round conversations at volume, scored consistently, available whenever the candidate is.
Answering, qualifying and routing calls that currently sit in a queue or go to voicemail.
Booking, confirming and rescheduling against a live calendar, including the follow-up call nobody makes.
Routine outbound checks — details confirmed, status updates collected, records updated as the call happens.
The same agent handling the languages your callers actually use, without a separate team per market.
Distressed callers, complex complaints, anything where being misunderstood carries real cost. Those calls belong to people.
Latency is a design constraint, not a metric
Every part of the pipeline is chosen against the clock: how fast speech is recognised, how quickly the model starts producing, how soon the first audio reaches the caller. Private hosting helps here — the inference sits next to the telephony rather than across a public API.
Call recordings and transcripts are personal data. They stay in the environment you approve.
- 01Speech in, fast
Streaming recognition tuned for your callers' accents and vocabulary, not a generic model.
- 02Turn-taking and barge-in
The agent stops when interrupted, handles silence, and does not talk over the caller.
- 03Grounded conversation
A defined script boundary with your systems behind it, so the agent can check and act mid-call.
- 04Handoff to a person
A live transfer with the transcript and context, on defined triggers — including the caller simply asking.
- 05Consent and retention
Disclosure at the start of the call, retention you set, storage in your region. How we deploy privately →
We built one for ourselves first
A voice agent that conducts screening interviews
- The problem
- We needed to screen candidates at volume without losing signal. First-round calls are structured and repetitive, but they still take a person's hour and the notes vary by interviewer.
- What we built
- A voice agent that runs the structured first round: asks the questions, follows up on thin answers, handles interruption and silence, and produces a consistent written assessment against the same criteria for every candidate.
- What we learned
- Most of the engineering went into the parts nobody demos — turn-taking, recovery when the line is bad, knowing when to stop and pass the call to a person. That is the work we bring to your build.
- Status
- Running in production as a product. Candidates speak to it; our team reads the assessments.
Next to this
Tell us about the calls nobody has time to make.
A discovery session is a working conversation, not a demo. We will tell you which of your calls should stay with people.