Ask a Hindi speaker to read the word कमल (lotus) and they will say kamal. Ask a naive speech synthesizer that converts each Devanagari letter to its textbook sound, and it will say kamala. The difference is a single vowel, but to a native listener it is the difference between a natural voice and one that sounds like it learned Hindi from a primer.
This vowel, and the rules for when it vanishes, is one of the central problems in Hindi text-to-speech. It sits alongside other challenges that the script and its romanizations only partly encode: retroflex versus dental consonants, aspiration, nasalization, and the enormous diversity of Indian personal and place names. For a voice agent that greets callers by name and confirms their addresses, these are not academic details. They are the first thing a caller notices.
The Inherent Vowel in Brahmic Scripts
Devanagari, Gujarati and most other scripts of South Asia are abugidas. In an abugida, each consonant letter carries a default vowel. In Devanagari, क is not just k; it is ka, where the vowel is a short central vowel usually transcribed as a and phonetically close to a schwa [ə]. To write a different vowel, you attach a vowel sign (मात्रा, mātrā): कि ki, का kā, कु ku. To write a consonant with no vowel at all, you use the virama (हलंत, halant): क् k, or you join the consonants into a conjunct, as in क्त kt.
On paper, then, the script looks perfectly explicit. Every consonant either has a vowel sign, a virama, or the inherent vowel. The problem is that ordinary Hindi spelling does not write the virama where the inherent vowel has stopped being pronounced. The spelling reflects an older, Sanskrit-like pronunciation, and modern speech has moved on.
Schwa Deletion in Hindi
Schwa deletion in Hindi has been described in detail by phonologists, notably Manjari Ohala in her work on Hindi phonology. The pattern has two main parts.
Word-final deletion
The inherent vowel at the end of a word is generally not pronounced:
- कमल: kamal, not kamala
- घर: ghar, not ghara
- राम: rām, not rāma
This is why so many Indian names appear in English with and without a final a. The Sanskrit form is Rāma; the everyday Hindi pronunciation is Rām.
Medial deletion
Inside a word, a schwa is typically deleted when it sits in a context like VC_CV: preceded by a vowel and a single consonant, and followed by a single consonant and a vowel.
- जनता: jantā, not janatā
- कमला: kamlā, not kamalā
- समझना: samajhnā, not samajhanā
Descriptions of the rule usually apply it scanning from the right end of the word, because deleting one schwa changes the context for its neighbours. In kamalā, the schwa after m has the vowel of ka and the consonant m on its left, and l plus ā on its right, so it deletes. In samajhanā, the schwa after jh has a + jh on its left and n + ā on its right, so it deletes; the schwa after m is then followed by jh and n, a consonant cluster, so it survives.
Why Simple Rules Fail
If the VC_CV rule were the whole story, grapheme-to-phoneme (G2P) conversion for Hindi would be easy. Researchers working on Hindi TTS have found that rule-based approaches get a large share of words right but leave a stubborn residue of errors, and those errors cluster in predictable places.
Morphology and compounds
The rule works well within a single morpheme but can fail across boundaries. Consider राजमहल (rājmahal, royal palace), a compound of rāj and mahal. Apply the rule mechanically from the right: the final schwa deletes, giving rājamahal. The schwa after m is now preceded by a vowel and m, and followed by h and a vowel, which matches VC_CV, so it deletes, giving rājamhal. The schwa after j is now before a cluster and survives. The output, rājamhal, is wrong. The correct pronunciation keeps each part intact: rāj + mahal. Getting this right requires knowing where the morpheme boundary is, and the script does not show it.
Inflection creates similar issues. Suffixes attach to stems in ways that can either trigger or block deletion, and a system that does not analyze the word's structure has to learn these patterns from data word by word.
Sanskrit-derived and formal vocabulary
Words borrowed more recently from Sanskrit, and formal or technical vocabulary, may retain schwas that the rule would delete, particularly in careful or formal speech. Pronunciation can vary between speakers, regions and speaking styles, so there is not always a single correct answer.
Loanwords and names
Words from Persian, Arabic, English and other languages, written in Devanagari, often do not follow native patterns. Personal names add a further layer: a name's pronunciation may follow family, regional or religious convention rather than the general phonology of Hindi.
Gujarati: Similar Problem, Different Details
Gujarati script is closely related to Devanagari, and Gujarati also tends to drop the inherent vowel in many contexts, including word-finally. કમળ (lotus) is pronounced roughly kamaḷ, without a final vowel. So the core challenge carries over.
The details, however, are not identical, and engineers should be cautious about reusing a Hindi G2P model. Some differences commonly noted by linguists of Gujarati:
- Vowel quality not shown in spelling. Gujarati distinguishes close-mid and open-mid vowels (roughly e versus ɛ, and o versus ɔ) in ways the script does not consistently mark.
- Breathy voiced vowels. Gujarati has murmured or breathy voice on vowels, often associated with a written h, and the phonetic realization is not a straightforward reading of the letters.
- Different vocabulary and morphology. Case markers, verb forms and common loanwords differ, so the contexts in which schwa deletion applies or is blocked also differ.
The honest summary is that Gujarati requires its own lexicon, its own rules or training data, and review by native speakers.
Beyond Schwa: Sounds English Voices Collapse
Even when the vowels are right, Hindi and Gujarati contain consonant contrasts that English does not have, and a synthesizer or listener trained on English tends to merge them.
Dental versus retroflex
Hindi distinguishes dental त t̪ and द d̪, made with the tongue against the teeth, from retroflex ट ʈ and ड ɖ, made with the tongue curled back. English /t/ and /d/ are alveolar, somewhere between the two, and Hindi speakers tend to hear them as retroflex. That is why English loanwords are usually written with ट and ड: टेबल for "table". A voice that uses English /t/ for a name containing त sounds noticeably foreign.
Aspiration
Hindi contrasts unaspirated and aspirated stops: क k versus ख kʰ, प p versus फ pʰ, and also breathy voiced stops such as भ bʱ and ध d̪ʱ. The pair काना kānā (one-eyed) and खाना khānā (food, to eat) differs only in aspiration. English has aspiration as an automatic variation, not a meaningful contrast, and English has no breathy voiced stops at all.
Nasalization
The anusvara ं and chandrabindu ँ both relate to nasal sounds but behave differently. Before a stop consonant, anusvara is usually pronounced as a nasal consonant that matches the stop's place of articulation: हिंदी hindī, संभव sambhav. Chandrabindu marks a nasalized vowel, as in हँसना hãsnā (to laugh). In careful spelling, हँस (laugh) and हंस (swan) are distinguished this way, although in everyday writing anusvara is often used for both, which leaves the G2P system to resolve the ambiguity.
Romanized Names: Where English TTS Breaks
Many business systems store customer names in Latin script: "Bhavesh Patel", "Dhruv Mehta", "Anjali Tripathi". These romanizations follow informal Indian conventions that look like English spelling but encode different sounds.
- "bh", "dh", "gh" usually represent breathy voiced stops. An English model tends to read "Bhavesh" with a plain b, and "Dhruv" with a plain d, losing the breathiness.
- "th" in "Tripathi" or "Mithun" usually represents an aspirated dental t̪ʰ, not the English θ in "think". English TTS often produces θ, which sounds clearly wrong.
- "t" and "d" can stand for either dental or retroflex consonants. "Ravindra" has a dental d; the romanization alone does not tell a system which one is meant in an unfamiliar name.
- "a" may represent short a, long ā, or a schwa that is not pronounced at all. "Ram" and "Rama" may be the same name.
- Stress is often placed by English rules, producing a rhythm that native speakers find unfamiliar.
The fundamental difficulty is that romanization is lossy. Going from "Dhruv" back to ध्रुव, and then to a pronunciation, is a reconstruction problem with real ambiguity. Knowing the likely language of origin of a name helps, as does context, such as the city or the language the caller is speaking.
Lexicons, Rules and Learned G2P
Production TTS systems typically combine several approaches, each with known trade-offs.
- Pronunciation lexicons give exact, verified pronunciations for known words. They are reliable but never complete, and names are the category least likely to be covered.
- Rule-based G2P encodes phonological knowledge such as schwa deletion. It is transparent and generalizes to unseen words, but breaks on exceptions and morphology.
- Learned G2P models, such as sequence-to-sequence neural models trained on lexicon data, can capture patterns rules miss. Their quality depends heavily on training data, and their errors can be harder to predict and correct.
- End-to-end TTS models learn pronunciation implicitly from paired text and audio. They can sound very natural but may still mispronounce rare words and names, and correcting a specific word can be less direct than editing a lexicon.
In practice, some mechanism for per-word overrides remains essential, because the words that matter most commercially are often the rarest ones in training data.
Why This Matters for Business
A voice agent that mispronounces a caller's name in its first sentence has already lost some trust. The same applies to locality names, product names and brand names. Place names such as Navrangpura or Vadodara, and surnames such as Chaudhary or Upadhyay, are everyday vocabulary for Indian callers but rare in generic training data.
The practical lessons for teams building Indian voice products are to treat pronunciation as a measurable quality dimension, to collect the names and places that actually appear in their customer base, to have native speakers listen to them, and to build a workflow for fixing errors quickly. At Cirio, pronunciation of names is something we consider part of basic courtesy on a call. The research literature gives a strong foundation, but the last mile is always specific: the names of the people and places your callers actually say.
Frequently Asked Questions
What is schwa deletion in Hindi?
Devanagari consonant letters carry an inherent short vowel, a schwa, unless another vowel sign or a virama is attached. In spoken Hindi this vowel is often not pronounced, especially at the end of a word and in certain medial positions, so कमल is said kamal rather than kamala. Speech synthesis systems must predict where the vowel disappears, because the spelling does not show it.
Why is schwa deletion hard for text-to-speech systems?
Simple rules capture much of the pattern but fail at morpheme boundaries, in compounds, in some Sanskrit-derived and borrowed words, and in names. Correct pronunciation often depends on how a word is built, which the script does not mark. Systems therefore combine rules, pronunciation dictionaries and learned models, and still need manual corrections.
Why do English TTS voices mispronounce Indian names?
Romanized Indian names use Latin letters in ways that differ from English spelling conventions. "th" in a name like Tripathi usually represents an aspirated dental stop, not the English th sound, and letters like t and d can stand for either dental or retroflex consonants. An English pronunciation model applies English rules and gets these wrong.
Does Gujarati have schwa deletion like Hindi?
Gujarati also tends to drop the inherent vowel in many of the same environments, such as word-finally, so the general challenge is similar. The details are not identical, and Gujarati has additional features its script does not fully mark, such as distinctions in mid vowel quality and breathy voiced vowels. A Hindi grapheme-to-phoneme model cannot simply be reused for Gujarati.