A speech synthesiser reads words. The text it receives, however, is full of things that are not words: ₹2,50,000, 05/06/2026, 3:45 PM, Dr., GST, 2-3 din, 9876543210. Before any audio can be produced, something has to decide how each of these should be spoken, and in which language. That step is called text normalisation (also verbalisation), and for Indian languages it is considerably harder than it first appears.
Getting it wrong is not a cosmetic issue. A voice agent that reads a loan amount as "two fifty thousand" or a delivery date in the wrong month has given the caller false information in a confident voice. This post walks through the main categories of the problem, the specific challenges of Hindi, Gujarati and Indian English, and the general approaches used to solve it.
What text normalisation does
In a typical text-to-speech pipeline, normalisation sits early: raw text comes in, and a sequence of speakable words comes out, which is then converted to phonemes and finally to audio.
The foundational framing comes from Richard Sproat and colleagues in their 2001 paper "Normalization of non-standard words," published in Computer Speech and Language. They proposed a taxonomy of non-standard words (NSWs) and argued that normalisation is fundamentally a classification problem followed by expansion. Given a token like 1995, the system must first decide what kind of thing it is (a year, a quantity, part of a phone number) and only then decide how to read it.
These categories are often called semiotic classes. A practical set for Indian business speech includes:
- Cardinal numbers: 45, 1,200
- Ordinals: 1st, 3rd, पहला
- Money: ₹499, Rs. 2.5 lakh
- Dates: 15/08/2026, 15 Aug
- Times: 10:30, 4 PM
- Telephone numbers and identifiers: mobile numbers, PIN codes, order IDs
- Measures: 5 kg, 2 BHK, 1,200 sq ft
- Abbreviations and acronyms: Dr., Pvt. Ltd., GST, UPI, EMI
- Ranges and fractions: 2-3, 1/2
The same surface string can belong to different classes. 12/10 may be a date, a fraction or a score. Deciding correctly usually depends on context.
The Indian numbering system
Indian English and Indian languages group large numbers differently from international convention. After the first three digits, grouping proceeds in twos:
- Indian: 2,50,000 read as "two lakh fifty thousand"
- International: 250,000 read as "two hundred fifty thousand"
The named units are hazaar (thousand, 10³), lakh (10⁵) and crore (10⁷). A normaliser for Indian speech has to handle several complications:
- Input may use either grouping. Text copied from an international system may say
250,000, while an invoice says2,50,000. Both should typically be spoken in lakhs for an Indian listener, and the normaliser should not treat commas as reliable signals of magnitude. - Decimal lakhs and crores.
2.5 lakhand1.25 croreare common in real estate and lending. In Hindi these are naturally read with fractional words: dhai lakh (ढाई लाख) for 2.5, sawa crore (सवा करोड़) for 1.25, dedh lakh (डेढ़ लाख) for 1.5, paune do lakh (पौने दो लाख) for 1.75. A literal reading like "do dashamlav paanch lakh" is correct but sounds unnatural.
Hindi numerals are not compositional
In English, once you know the words for 1 to 19 and the tens, you can generate every number up to 99 by rule: forty plus five is forty-five. Hindi does not work this way. Each number from 1 to 100 has its own word, and many of them are not predictable from their parts:
- 29: उनतीस (unatees)
- 45: पैंतालीस (paintaalees)
- 49: उनचास (unchaas)
- 68: अड़सठ (adsath)
- 99: निन्यानवे (ninyaanve)
Numbers such as 29, 39 and 49 are formed as "one less than" the next ten (the prefix un-), and the stem for each unit digit changes with the tens. Gujarati has a similarly irregular set; for example 25 is પચ્ચીસ (pachchees). For this reason, practical normalisers store the words for 1 to 99 in a lookup table per language and apply rules only above that level, composing with sau (hundred), hazaar, lakh and crore: 3,45,678 becomes teen lakh paintaalees hazaar chhah sau athhattar.
Currency, dates and times
Rupees and paise
₹1,499 is ek hazaar chaar sau ninyaanve rupaye in Hindi and "one thousand four hundred ninety-nine rupees" in English. Decimals need care: ₹249.50 should be read as rupees and paise (do sau unchaas rupaye pachaas paise), not as a decimal number. Prefixes such as Rs., INR and ₹ all map to the same reading, and the symbol can appear before or after the number.
Dates
The main hazard is ambiguity between day-first and month-first formats. In India, 05/06/2026 almost always means 5 June, but text originating from systems configured for US conventions means 6 May. A normaliser cannot resolve this from the string alone; it needs to know the source convention, and where possible systems should pass dates in an unambiguous structured form rather than as formatted text.
Reading styles also vary by language. Hindi typically says paanch June or paanch tareekh, and years are commonly read as do hazaar chhabbees for 2026, while 1998 is often unnees sau atthaanve. English speakers in India may say "fifth June" or "June fifth." Choosing one style and applying it consistently matters more than any particular choice.
Times
3:30 in Hindi is naturally saadhe teen baje, 2:15 is sawa do baje and 2:45 is paune teen baje. A literal "teen baj kar tees minute" is understandable but stilted. AM and PM are frequently rendered as subah, dopahar, shaam or raat according to the hour, which requires a small mapping rather than a direct translation.
Phone numbers, PIN codes and identifiers
Identifiers must never be read as quantities. 9876543210 read as a cardinal would be "nine hundred eighty-seven crore..." which is useless to a caller trying to write it down.
- Mobile numbers are read digit by digit, usually with grouping. Two groups of five is common for ten-digit numbers, though some listeners prefer other patterns. A leading
+91is often read as "plus nine one" or dropped when context makes it clear. - PIN codes are six digits and are generally read digit by digit (chaar, zero, zero, zero, zero, one for 400001), sometimes in groups of three.
- Order and booking IDs mix letters and digits and should be spelled out with clear pauses. Repeated digits are sometimes read as "double five" in Indian English, a convention worth supporting because listeners expect it.
Ordinals, abbreviations and ranges
Ordinals in Hindi are irregular for the first few values (पहला pehla, दूसरा doosra, तीसरा teesra, चौथा chautha, छठा chhatha) and then largely regular with the suffix -vaan (पाँचवाँ paanchvaan). They also inflect for gender: teesri manzil (third floor) versus teesra din (third day). Getting gender agreement right requires knowing the noun that follows.
Abbreviations fall into different reading types:
- Expanded to a word:
Dr.as "doctor,"Pvt. Ltd.as "private limited,"St.as "street" or "saint" depending on context. - Read as letters:
GSTas "G S T,"UPIas "U P I,"EMIas "E M I,"PANoften as "pan" and sometimes spelled. - Read as words: some acronyms are conventionally pronounced as words, which must be listed explicitly.
Ranges like 2-3 din should become do se teen din in Hindi or "two to three days" in English. The hyphen must not be read as "minus," and it must not be confused with a date separator or a phone number group.
The code-mixing question
Indian business text is frequently mixed. Consider: "Your EMI of ₹12,500 is due on 5th October." If the voice is speaking Hindi, or the caller has been speaking Hindi throughout, should the amount be baarah hazaar paanch sau or "twelve thousand five hundred"?
There is no universal answer, and good systems treat it as a policy decision:
- Follow the matrix language. Numbers take the language of the surrounding sentence. This is consistent but can sound odd when a single English sentence appears in a Hindi conversation.
- Follow the caller. Many urban Hindi speakers say amounts and dates in English even when speaking Hindi ("aapka EMI twelve thousand five hundred hai"). Mirroring this can feel more natural.
- Prefer comprehension for critical values. For amounts, OTPs and dates, the reading the listener will understand most reliably is what matters. For some audiences that is English numbers; for others, especially older callers or those in smaller towns, it is Hindi or Gujarati.
Whatever the policy, the normaliser needs to know which language the output voice will speak and must produce words in a script and phonetic form that the synthesiser can pronounce correctly.
Rule-based and neural approaches
Traditionally, normalisation has been built from hand-written grammars, often compiled into weighted finite-state transducers (WFSTs). The Kestrel system, described by Peter Ebden and Richard Sproat in 2015, is a well-known example of this style. The advantages are predictability and auditability: a rule for rupee amounts either fires or does not, and when it fires, the output is exactly what the grammar specifies. The disadvantage is effort. Every language, class and edge case must be written and tested by hand.
Neural normalisation treats the task as sequence-to-sequence transduction, learning from pairs of written and spoken text. It handles context and unusual inputs more flexibly. The risk, highlighted by Sproat and Navdeep Jaitly in their work on neural text normalisation, is what they called unrecoverable errors: outputs that read fluently but change the meaning, such as speaking a different number from the one written. In a voice agent quoting a price or a date, that is the worst possible failure, because the listener has no way to notice it.
This is why many production systems use hybrid designs: learned models for classification or ambiguous context, with deterministic expansion or explicit verification for high-stakes classes such as money, dates, times and identifiers. The same caution applies when a general-purpose language model is asked to write out numbers as words: it is convenient, but its output should be checked rather than trusted.
Building it well
Text normalisation rarely gets attention until it fails on a live call. For teams building voice products for Indian audiences, a few practices go a long way:
- Build a test set of real sentences from your domain, covering amounts, dates, identifiers and mixed-language text, with reference readings approved by native speakers.
- Keep structured values structured for as long as possible, so a date is still a date, not a string, when it reaches the normaliser.
- Treat language choice for numbers as an explicit product decision.
At Cirio, we consider this unglamorous layer part of what makes a voice agent trustworthy. A caller may forgive a slightly robotic voice; they will not forgive being told the wrong amount.
Frequently Asked Questions
What is text normalisation in text to speech?
Text normalisation is the step that converts written forms such as numbers, currency, dates, abbreviations and symbols into the words a speaker would actually say. For example, ₹1,500 becomes ek hazaar paanch sau rupaye in Hindi. Without it, a speech synthesiser either skips these tokens or reads them incorrectly.
Why are Hindi numbers hard for speech synthesis?
Hindi number words from 1 to 100 are largely irregular, so 45 is पैंतालीस (paintaalees) and 49 is उनचास (unchaas), not a simple combination of forty and five. They cannot be generated by a short rule and are normally stored in a lookup table. Larger numbers then combine these words with sau, hazaar, lakh and crore.
How should Indian phone numbers be read by a voice agent?
Phone numbers should be read as digit sequences, never as a single large number. A ten-digit mobile number is commonly grouped, often in two groups of five, with short pauses, so the listener can write it down. The words for the digits should match the language of the surrounding sentence.
Is neural text normalisation better than rule-based normalisation?
Neural models handle context and unusual inputs well, but they can occasionally produce fluent yet wrong output, such as reading a different number from the one written. Rule-based grammars are predictable but costly to extend. Many practical systems combine the two, with rules or checks guarding high-stakes categories like amounts and dates.