Word error rate (WER) is the lingua franca of speech recognition. Research papers report it, benchmarks rank systems by it, and procurement discussions often reduce to comparing two WER figures. It is a useful metric, and anyone evaluating speech systems should understand it thoroughly.
It is also easy to misread. A recognizer with a respectable WER can still fail on the one word in a sentence that matters, and a recognizer that looks worse on paper may serve users better. For Indian languages, telephone audio and code-mixed speech, the gap between a WER number and real-world usefulness can be especially wide. This post explains why, and what to measure instead of, or alongside, a single headline figure.
What WER Actually Measures
WER compares a recognizer's output (the hypothesis) with a human-produced reference transcript. The two word sequences are aligned using minimum edit distance, and three types of errors are counted:
- Substitutions (S): a reference word replaced by a different word.
- Deletions (D): a reference word missing from the hypothesis.
- Insertions (I): an extra word in the hypothesis.
With N as the number of words in the reference:
WER = (S + D + I) / NSuppose the reference is "mera order number teen chaar paanch hai" (7 words) and the hypothesis is "mera order number teen char paanch". "char" for "chaar" is a substitution and the missing "hai" is a deletion, so WER is 2/7, about 28.6 percent. Note that one of these "errors" is merely a spelling variant.
Why WER can exceed 100 percent
Because insertions add to the numerator while the denominator counts only reference words, WER has no upper bound. If a caller says "haan" and the recognizer transcribes background television as eight words, WER for that utterance is 800 percent. In aggregate reporting this is rare, but on short utterances, which dominate phone conversations, it is a real effect.
WER is also not a true percentage of words that were wrong. Morris, Maier and Green (2004) proposed alternatives such as match error rate and word information lost that are bounded between 0 and 1, partly for this reason. WER remains the convention, but it is worth remembering what it is and is not.
Normalization: The Hidden Variable
Before computing WER, both reference and hypothesis are usually normalized. The choices made here can change results more than the difference between two competent systems.
Numbers, punctuation and casing
Is "25" equivalent to "twenty five" or "पच्चीस"? Is "Rs. 500" the same as "five hundred rupees"? Should "OK", "okay" and "ok" match? If the reference uses digits and the recognizer writes words, a perfect transcription can score poorly. Most evaluations lowercase text and strip punctuation, and many apply rules for numbers, but these rules are rarely identical between studies.
Spelling variants in Hindi
Devanagari Hindi has several widely accepted spelling alternations:
- Chandrabindu versus anusvara: हाँ and हां (haan).
- Forms such as गए and गये (gaye), or लिए and लिये (liye).
- Presence or absence of nukta in loanwords, and variant spellings of English loanwords such as ऑर्डर and आर्डर (order).
A human would consider each pair correct. Unnormalized WER counts each mismatch as a full substitution. On conversational Hindi, where such words are frequent, this can inflate WER considerably.
Script choice
Code-mixed Indian speech raises a more fundamental question: which script should the reference use? A caller who says "mera delivery address change karna hai" could be transcribed as:
- मेरा डिलीवरी एड्रेस चेंज करना है (all Devanagari),
- मेरा delivery address change करना है (mixed script), or
- mera delivery address change karna hai (all Roman).
All three are reasonable. If the reference follows one convention and the recognizer another, WER will be extremely high even though the recognition is flawless. Any evaluation must fix a convention, document it, and either transliterate outputs to match or score in a script-agnostic way.
Character Error Rate and Related Metrics
Character error rate (CER) applies the same formula at the character level. It is standard for languages written without spaces, and it is helpful for Indic languages too.
Gujarati and Hindi attach postpositions and inflections in ways that create long or variably segmented words. In Gujarati, "ઘરમાં" (gharmaan, "in the house") is written as one word; a recognizer that outputs "ઘર માં" with a space makes one substitution and one insertion at word level, a 200 percent error on a one-word reference, while CER registers only a single inserted space. CER therefore helps separate segmentation and spelling issues from genuine misrecognitions.
For Indic scripts, one caveat applies: "character" should be defined carefully. Scoring at the level of Unicode code points counts a matra error as one character, while scoring at the grapheme cluster level treats the whole visual syllable as the unit. Either is defensible, but the choice must be consistent.
For code-mixed speech, researchers working on Mandarin and English have used a mixed error rate, which scores Mandarin at the character level and English at the word level. Analogous hybrid scoring can be useful for Hindi and English mixing, together with reporting error rates separately for words in each language, so that weakness on English loanwords or on native vocabulary is visible rather than averaged away.
Entity and Slot Accuracy
In business conversations, a small number of words carry most of the value. Mishearing "kal" as "kaal" is harmless; mishearing a phone number digit, a customer's name or an amount is not.
Consider a caller saying "नौ आठ दो पाँच, शून्य शून्य एक, तीन चार सात" (a ten-digit mobile number read in groups). Out of a transcript of forty words from the whole turn, one wrong digit contributes a tiny amount to WER. For the business, the entire callback fails.
Useful entity-level measures include:
- Exact match rate for structured values such as phone numbers, PIN codes, order IDs, dates and amounts, after normalizing to a canonical form (for example, converting spoken Hindi or English digits into a digit string).
- Name accuracy, ideally with phonetic or transliteration-aware matching, since "Patel", "पटेल" and "પટેલ" are the same name.
- Slot precision and recall when the application extracts fields such as appointment time, product and quantity.
These metrics require the test set to be annotated with entities, which is more work than plain transcription, but they measure what actually determines whether a call succeeds.
Semantic Error and Task Success
Beyond individual entities, one can ask whether the transcript preserves meaning. Semantic error rate approaches judge whether the hypothesis would lead to the same interpretation as the reference, for instance by checking if a language understanding component extracts the same intent and slots from both. More recent work has also explored using learned semantic similarity models to score transcripts, though such scores inherit the biases of the model used and should be validated against human judgement.
The most direct measure is task success: did the end-to-end system book the right appointment, capture the correct lead, or route the call correctly? Task success depends on much more than recognition, so it cannot isolate the recognizer's contribution on its own. Combined with the metrics above, though, it closes the loop between technical accuracy and outcomes. A recognizer change that lowers WER but does not improve, or even reduces, task success deserves close scrutiny.
Building a Representative Test Set
No metric is better than the data it is computed on. Public benchmarks are valuable for research, but they are often read or broadcast speech recorded at wideband sampling rates. Real phone calls differ in almost every respect.
A test set for Indian telephony applications should reflect:
- Channel. Traditional telephone audio is narrowband, typically sampled at 8 kHz, with codecs, compression and packet loss. Evaluating on wideband audio and deploying on phone lines will overestimate accuracy.
- Speakers and accents. Hindi spoken in Ahmedabad, Lucknow and Bengaluru sounds different. Include a spread of regions, ages and genders proportional to the real caller base.
- Environment. Traffic, shop floors, television, speakerphone use and people talking in the background.
- Conversational style. Short answers, hesitations, self-corrections, backchannels such as "haan ji", and code-mixing, rather than carefully read sentences.
- Domain vocabulary. Product names, local place names and the specific entities the application needs to capture.
Transcription guidelines matter as much as audio selection. Annotators need explicit conventions for script, numbers, spelling variants, disfluencies and unintelligible segments, and ideally a subset should be double-transcribed to measure inter-annotator agreement. If two careful humans disagree on a significant fraction of words, no recognizer can meaningfully score below that level.
Finally, keep the test set strictly separate from any data used for training or tuning, and refresh it periodically as callers and use cases evolve.
Significance with Small Test Sets
Domain-specific test sets are often small: a few hundred utterances, perhaps a few hours of audio. Differences of a point or two in WER on such sets may be noise.
Errors are also not independent across words. Utterances from the same speaker or the same noisy call tend to fail together. Two established approaches account for this:
- Bootstrap resampling at the utterance or speaker level, as described for ASR by Bisani and Ney (2004), produces confidence intervals for WER and for the difference between two systems.
- Matched-pairs tests, such as the matched pairs sentence-segment word error test associated with Gillick and Cox (1989) and implemented in standard NIST scoring tools, compare two systems on the same segments.
When confidence intervals overlap substantially, the honest conclusion is that the systems perform similarly on this data, and other factors should decide.
A Practical Evaluation Mindset
WER is a sound starting point, not a destination. A useful evaluation for real deployments combines several layers: normalized WER and CER with documented conventions, error rates broken down by language in code-mixed speech, exact-match accuracy on the entities that matter, and end-to-end task success, all measured on audio that genuinely resembles production calls and reported with confidence intervals.
This takes more effort than running a single script against a public benchmark, but it answers the question that actually matters: will this system understand the people who call? At Cirio we consider that question central, and we encourage any business evaluating voice technology to ask vendors not only for a WER figure, but for how it was measured, on what audio, and what happens to the phone numbers.
Frequently Asked Questions
How is word error rate calculated?
Word error rate is the minimum number of substitutions, deletions and insertions needed to turn the recognizer's output into the reference transcript, divided by the number of words in the reference. The alignment is computed with edit distance, and the result is usually reported as a percentage.
Can word error rate be higher than 100 percent?
Yes. Because insertions are counted but the denominator is only the number of reference words, a recognizer that outputs many extra words can exceed 100 percent. This often happens when noise or hold music is transcribed as speech.
When should character error rate be used instead of WER?
Character error rate is useful when word boundaries are ambiguous or when words are long and heavily inflected, because a single wrong suffix does not count as an entire word error. For Indic scripts it is often reported alongside WER to separate spelling-level mistakes from genuine recognition failures.
What is the best way to evaluate speech recognition for a business use case?
Build a test set from audio that matches real conditions, such as narrowband phone calls with the accents and noise your callers have. Then measure what the application needs, such as accuracy on phone numbers, names and amounts, and whether the downstream task succeeded, in addition to normalized WER.