A voice agent that conducts screening interviews
We needed to screen candidates at volume without losing signal, so we built it for ourselves. It runs in production on our own hiring, and it is why we can talk about voice agents from experience rather than a datasheet.
- Client
- Ourselves
- Service
- Voice AI Agents
- Product
- Vaani
- Status
- In production on our own hiring
Hiring engineers means a lot of first-round calls. They are structured, they cover the same ground, and they are necessary — but each one takes an interviewer's hour, and scheduling them takes longer than holding them.
Two problems compounded. Volume meant the calls competed with client work, so candidates waited. And because different people ran them, the notes varied: one interviewer probed a thin answer, another moved on, and comparing two candidates meant comparing two different conversations.
Screening at scale without losing signal is exactly the shape of problem we tell clients to bring us. It seemed reasonable to solve it on ourselves first.
A voice agent that runs the structured first round: asks the questions, follows up where an answer is thin, and produces a written assessment against the same criteria for every candidate.
Not a form read aloud. The agent follows up on a vague answer, handles a candidate thinking out loud, and copes with someone talking over it.
Scored against criteria set for the role, with the transcript attached so any judgment can be checked rather than taken on trust.
The call happens when it suits them rather than when a diary opens, which removed the scheduling delay entirely.
Vaani rejects nobody. It produces the first pass; the hiring decision stays with people, which is also what the EU AI Act requires of recruitment systems.
Most of the engineering went into the parts nobody demonstrates.
If the reply does not begin quickly enough, the candidate assumes the line has dropped and starts talking again. Every component gets chosen against the clock.
The agent has to stop mid-sentence and listen. An agent that talks over people is worse than a form.
Accents, background noise, patchy mobile connections. Recovery behaviour matters more than accuracy on a clean recording.
A confident wrong answer is worse on a call than in text, because there is no time to check it. Defined limits and a clean handoff are part of the build, not an afterthought.
Under the EU AI Act, AI used in recruitment sits in a high-risk category, and call recordings are personal data from the first second.
That shaped the build: disclosure to the candidate at the start of the call, no automated rejection, transcript and scoring retained together so a judgment can be reviewed, and processing on private models with residency and retention set by whoever deploys it.
How we deploy privately →Voice fails in production, not in the demo
Every voice agent sounds good in a scripted demonstration. The difference shows on a bad line, with an interruption, at the fiftieth call of the day. We have already had those calls fail and fixed them.
This is the only case study on the site where the client is us, and it earns its place for that reason: it is the proof that we put a voice agent into production, not a slide about one.
See Vaani →Hear it before you decide anything.
A walkthrough is a live call with the agent, using your own role and criteria. Twenty minutes, and you will know whether it holds up.