Testing traditional software is comparatively tidy: given an input, check the output. A conversational voice agent breaks that model. The caller can say anything, in any order, in Hindi, Gujarati, English or a mix, over a noisy 8 kHz phone line, while interrupting. Two perfectly acceptable conversations for the same task may share almost no words. And many of the qualities that matter most, such as whether the agent sounded natural or whether the caller felt understood, are hard to reduce to a number.
Yet without measurement, every change to a voice agent is a guess. A prompt edit that fixes one scenario can quietly break three others. This post sets out a general framework for evaluating voice agents: what to measure, how to measure it at scale, and where each method can mislead you.
Why conversation is hard to evaluate
Several properties make voice agents harder to evaluate than most systems:
- No single reference answer. For a given caller turn, many responses are acceptable. Metrics based on overlap with a reference response, borrowed from machine translation, correlate poorly with human judgement in open dialogue.
- Compounding errors. A recognition error in turn two can derail everything that follows. Evaluating turns in isolation misses these cascades.
- Interaction effects. The caller's next utterance depends on what the agent said. A fixed script of caller lines cannot test a system whose behaviour changes the conversation's path.
- Distribution shift. Real callers use accents, vocabulary, background environments and code-mixing patterns that test data rarely captures fully.
The practical answer is to evaluate at several levels at once and to accept that each level answers a different question.
Component, turn-level and end-to-end metrics
Component metrics
Each stage of a voice system has established measures:
- Speech recognition is commonly measured with word error rate (WER). For Indian languages, character error rate is often reported too, since word segmentation and spelling variation make WER noisy. Code-mixed speech adds a further problem: "booking" may be transcribed in Roman script or as बुकिंग, and a naive WER calculation counts that as an error even though the meaning is identical. Normalising script and spelling before scoring is essential.
- Speech synthesis is traditionally evaluated with human listening tests such as mean opinion score (MOS), plus targeted checks for pronunciation of names, numbers and domain terms.
Component metrics are fast, reproducible and useful for diagnosis. Their limitation is that they are not the outcome. A lower WER overall may still come with worse recognition of the exact words that matter, like phone numbers.
End-to-end metrics
End-to-end metrics look at the conversation as the caller experiences it:
- Task completion rate: did the call achieve its purpose, such as a booked appointment, a qualified lead or a resolved query?
- Escalation and abandonment: how often did the caller ask for a human, hang up early or go silent?
- Correctness of outcomes: was the data captured accurate? A "completed" booking with the wrong date is worse than no booking.
Turn-level quality
Between components and whole conversations sits the individual turn. Useful questions to ask of each agent turn include:
- Was it relevant to what the caller said?
- Was it factually grounded in the business information provided, without inventing prices, policies or availability?
- Did it follow instructions, such as staying within scope and respecting compliance wording?
- Was it appropriate for speech: short enough to follow by ear, free of formatting artefacts, and in the caller's language and register?
- Did it handle turn-taking correctly, without talking over the caller or leaving an awkward silence?
Turn-level labels are especially useful for locating failures. When a conversation fails, identifying the first bad turn usually points directly at the cause.
Measuring latency properly
Responsiveness shapes how natural an agent feels, and it is widely measured badly.
Report percentiles, not averages. Latency distributions in systems that chain several networked stages are typically right-skewed. The mean can look healthy while a meaningful fraction of turns are slow. Reporting the median (p50) alongside p95 and p99 reveals the tail. Because a single call contains many turns, the tail matters more than intuition suggests: if one turn in twenty is slow, roughly four in ten calls of ten turns will contain at least one slow turn, assuming turns are independent.
Define the interval precisely. "Latency" can mean many things. The measure closest to caller experience is usually the time from when the caller stops speaking to when they hear the start of the agent's reply. Internal stage timings are valuable for diagnosis, but they should add up to, and be reconciled with, the caller-perceived interval.
Evaluating at scale
Simulated callers
Real calls are the gold standard, but they arrive slowly and cannot be replayed against a new version of the agent. Simulated callers fill this gap. A common approach uses a language model given a persona and a goal, for example: "You are a retired schoolteacher in Ahmedabad who speaks mostly Gujarati, wants to reschedule a clinic appointment, and is unsure of the exact date." The simulator then holds a full conversation with the agent.
Well-designed simulations vary several dimensions:
- Goals, including ones the agent cannot fulfil.
- Behaviour: cooperative, impatient, confused, verbose, or deliberately off-topic.
- Language: pure Hindi, pure English, Hinglish, Gujarati, switching mid-call.
- Adversarial inputs: attempts to extract information the agent should not share, or to make it ignore its instructions.
Simulations have real limits. Language models acting as callers tend to be more articulate, more cooperative and more consistent than real people. They rarely produce the hesitations, self-corrections, background interruptions and half-sentences that fill real phone calls. When run as text, they skip recognition errors entirely. Results from simulation should be read as evidence about logic and instruction-following, not as a prediction of real-world success rates.
LLM-as-judge and its biases
Human review does not scale to thousands of conversations, so many teams use a language model to grade transcripts against a rubric. Good rubrics are specific and observable: "Did the agent confirm the appointment date back to the caller before ending the call? Yes or no" is far more reliable than "Rate the conversation quality from 1 to 10."
The research literature has documented systematic biases in model judges. Zheng and colleagues, in their 2023 study of language models as judges, described several:
- Position bias: when comparing two responses, a judge may favour whichever appears first (or second), regardless of quality.
- Verbosity bias: longer, more detailed answers tend to be rated higher, which is exactly the wrong preference for voice, where brevity matters.
- Self-enhancement bias: a judge may favour outputs resembling its own style.
Common mitigations include swapping the order in pairwise comparisons and averaging, asking for binary or narrowly scaled judgements, requiring the judge to cite the specific turn that supports its verdict, and explicitly instructing that shorter answers are preferred when equally correct.
Calibrating against humans. A judge is only useful if it agrees with the people whose judgement you care about. The standard practice is to have human reviewers label a representative sample, run the judge on the same sample, and measure agreement using statistics such as Cohen's kappa, which corrects for chance agreement. Where agreement is weak on a criterion, refine the rubric or keep that criterion under human review. Calibration should be repeated when the rubric, the judge or the agent changes significantly.
Building test suites
Regression suites from real failures
The most valuable test cases are the ones that already went wrong. Each time a real call reveals a failure, such as a misheard PIN code, a caller who switched to Gujarati and lost the agent, or an agent that promised a discount it could not give, it should become a permanent test case with a clear pass condition.
Over time this produces a regression suite that reflects the actual risks of your deployment rather than imagined ones. Run it on every change to prompts, models, configuration or knowledge content. Track results per case, not only in aggregate, so that a single important regression is not hidden by improvements elsewhere.
Testing at the audio level
Text-based tests skip the hardest part of a voice system. Audio-level testing sends real or synthesised speech through the full pipeline, including telephony conditions:
- Narrowband audio: many phone calls are carried at 8 kHz, which removes high-frequency cues that help distinguish sounds such as s and f. Test audio should be downsampled and passed through a representative codec.
- Noise: traffic, crowded markets, television, fans, wind on a two-wheeler.
- Accents and dialects: speakers from different regions, and elderly speakers.
- Code-mixing: sentences that switch language mid-phrase, the norm for many urban callers.
- Critical content: numbers, names, addresses and dates, where a single recognition error changes the outcome.
- Turn-taking behaviour: interruptions, backchannels and long pauses, which only exist in audio.
Recorded human speech is preferable for realism; synthesised speech is convenient for coverage but tends to be cleaner than real callers.
Learning from production
A/B testing
Offline evaluation reduces risk; only production tells you the effect on real callers. Controlled experiments route a portion of traffic to a new version and compare outcomes such as task completion, escalation, call duration and post-call feedback.
Some principles apply particularly to voice:
- Choose outcome metrics before the test and avoid checking repeatedly for significance, which inflates false positives.
- Allow enough volume. Differences in completion rates are often small, and call volumes in a single deployment may need weeks to detect them reliably.
- Watch guardrail metrics such as complaints, compliance issues and error rates, and stop an experiment early if they degrade.
Reviewing recordings responsibly
Listening to real calls remains irreplaceable, but recordings and transcripts contain personal information: names, phone numbers, addresses, account details and sometimes health or financial information. Responsible review practices include redacting personal data from transcripts before broad access, restricting raw audio to a small group of authorised reviewers, sampling rather than bulk exporting, honouring retention limits and consent obligations, and keeping an audit trail of who accessed what. For businesses operating in India, these practices should also be reviewed against obligations under the Digital Personal Data Protection Act, 2023.
Bringing it together
No single method is enough. Component metrics diagnose, end-to-end metrics decide, simulations and judges provide scale, human review provides truth, and production experiments confirm impact. The discipline lies in connecting them: every real failure becomes a test, every automated score is checked against people, and every release is measured before it is trusted.
At Cirio, we think of evaluation as part of the product rather than an afterthought, because callers only ever experience the conversation, never the benchmark.
Frequently Asked Questions
How do you measure the quality of a voice agent?
Quality is measured at several levels: component metrics such as recognition accuracy and synthesis quality, turn-level measures such as response relevance and latency, and conversation-level outcomes such as task completion and escalation rate. No single number is sufficient, because a system can score well on components while failing callers end to end.
Why use latency percentiles instead of averages?
Averages hide the slow responses that callers actually notice. A system with a good mean may still have one turn in twenty that takes several times longer, and over a ten-turn call that tail is easily hit. Reporting p50, p95 and p99 shows both typical and worst-case experience.
Can a language model be used to grade voice agent conversations?
Yes, using a language model as a judge with a clear rubric is a practical way to score large numbers of transcripts. However, such judges have known biases, including sensitivity to position, preference for longer answers and a tendency to favour outputs similar to their own. Their scores should be calibrated against human labels before being trusted.
What are simulated callers in voice agent testing?
Simulated callers are automated counterparts, often driven by a language model with a persona and goal, that hold full conversations with the agent under test. They make it possible to run hundreds of scenarios before release. They do not fully reproduce real human behaviour, so they complement rather than replace real call review.