Back to Blog
NLP9 min read1,715 words

Token Fertility in Indic Languages: Why Hindi Requires More Tokens Than English

How subword tokenizers treat Devanagari and Gujarati text, and why it matters for language models in India.

Abstract cover showing a line of Devanagari characters breaking apart into many small coloured fragments

Language models do not read letters or words. They read tokens, the units produced by a tokenizer before any neural computation begins. For English, a token is often a whole common word or a large fragment of one. For Hindi, Gujarati and many other Indian languages, the same tokenizer may break a single word into several small pieces, sometimes down to individual bytes.

This difference is easy to overlook because it is invisible in a chat window. It nonetheless shapes how much text fits in a model's context, how long generation takes, and in some cases how well the model understands the language at all. For anyone building language technology for Indian users, it is worth understanding precisely where the imbalance comes from.

What Subword Tokenizers Do

Early neural language systems used fixed word vocabularies, which failed on rare words, names and inflected forms. Subword tokenization solved this by representing text with a vocabulary of frequent fragments that can be combined to spell anything.

Byte Pair Encoding

Byte Pair Encoding (BPE), adapted for neural machine translation by Sennrich, Haddow and Birch (2016), starts from individual characters and repeatedly merges the most frequent adjacent pair in the training corpus into a new symbol. After many merges, common words become single tokens, while rare words are spelled out from smaller pieces.

WordPiece and Unigram

WordPiece, described by Schuster and Nakajima (2012), is similar but chooses merges that most increase the likelihood of the training data. The unigram language model approach (Kudo, 2018) works in the opposite direction: it starts with a large candidate vocabulary and prunes pieces that contribute least to corpus likelihood.

SentencePiece and byte-level fallback

SentencePiece (Kudo and Richardson, 2018) implements BPE and unigram directly on raw text, treating whitespace as an ordinary symbol, which makes it convenient for languages without clean word boundaries. Many modern tokenizers also operate at the byte level or include byte fallback: when a character or sequence is not in the vocabulary, it is represented by its underlying UTF-8 bytes. This guarantees that any text can be encoded, but unfamiliar scripts pay a steep price for that guarantee.

Fertility: A Simple, Revealing Metric

Rust et al. (2021), in "How Good is Your Tokenizer?", used fertility (the average number of subword tokens per word) as a way to compare tokenizers across languages, and found that tokenizer quality for a language is associated with downstream model performance in that language. Petrov et al. (2023), in "Language Model Tokenizers Introduce Unfairness Between Languages", showed that the same content can require very different token counts depending on the language it is written in, with many non-Latin-script languages among the most affected.

Fertility is not a perfect measure. Word boundaries are defined differently across languages, and a highly agglutinative language will naturally have longer words. Still, when a translation of the same sentence consistently produces several times as many tokens in one language as in another, the tokenizer is clearly representing the two languages with very different efficiency.

Where the Imbalance Comes From

Corpus skew

A BPE or unigram vocabulary is a compressed summary of its training data. If that data is overwhelmingly English, the vocabulary budget is spent on English words and fragments: "ing", "tion", "customer", "delivery". Hindi sequences such as "करना" (karna) or "चाहिए" (chahiye) may be common in Hindi text, but if Hindi is a small fraction of the corpus they may never be frequent enough, in absolute terms, to be merged into single tokens. Gujarati, with a smaller share of web text than Hindi, is typically affected even more.

UTF-8 byte lengths

The encoding itself contributes. In UTF-8, ASCII characters occupy one byte each. Code points in the Devanagari block (U+0900 to U+097F) and the Gujarati block (U+0A80 to U+0AFF) occupy three bytes each.

Take the greeting namaste:

  • English-style hello is 5 code points and 5 bytes.
  • नमस्ते is 6 code points: न, म, स, the virama ्, त, and the vowel sign े. That is 18 bytes.
  • નમસ્તે in Gujarati has the same structure: 6 code points, 18 bytes.

If a byte-level tokenizer has learned no merges covering these sequences, a single short Hindi word could in principle become 18 tokens, while an English word of similar length becomes one or two.

Matras, viramas and conjuncts

Indic scripts are abugidas. A consonant carries an inherent vowel, other vowels are written as dependent signs called matras, and consonant clusters are formed with a virama, often rendered as a visual ligature. What a reader perceives as one or two characters can be several code points.

The word क्षेत्र (kshetra, "area" or "field") appears to a reader as a few visual units, but it consists of 7 code points (क, ्, ष, े, त, ्, र) and 21 bytes. A tokenizer that does not respect these units may split a conjunct in the middle, separating a consonant from its virama, producing fragments that carry little linguistic meaning on their own.

Spelling variation adds further fragmentation. हाँ (with chandrabindu) and हां (with anusvara) are both common ways to write haan. They differ in one code point, so a tokenizer sees them as different sequences, splitting the frequency of the word across two forms.

A Hypothetical Illustration

Consider a customer message and its English equivalent:

  • "मेरा ऑर्डर अभी तक नहीं आया है"
  • "My order has not arrived yet"

Imagine an English-heavy tokenizer. For the English sentence, every word is common, so it might produce about one token per word. For the Hindi sentence, common short words like है or नहीं might have their own tokens if Hindi was reasonably represented, but a transliterated loanword such as ऑर्डर might split into several pieces, and less common forms could fall back to bytes. The Hindi sentence could plausibly require two to four times as many tokens as the English one. This is an illustration of the pattern, not a measurement of any specific tokenizer; the actual ratio depends heavily on the vocabulary and its training data.

Romanized Hindi: A Different Trade-Off

Much Hindi and Gujarati on Indian phones is typed in Roman script: "mera order abhi tak nahi aaya hai". Because this reuses Latin characters and many English-like fragments, it often tokenizes more compactly than Devanagari.

The trade-off is consistency. The same word appears as nahi, nahin, nai or nhi; "hai" may be "h" in casual typing. A model has to learn that all of these are equivalent, and romanized Indic text is less standardised than either English or native-script Hindi in most corpora. Romanization can also be ambiguous: "kal" means both "yesterday" and "tomorrow" in Hindi even in native script, but romanized text loses additional distinctions such as retroflex versus dental consonants (ट versus त) and vowel length.

For speech applications the question arises naturally, because transcripts can be produced in either script. Choosing a script is a system-level design decision that affects recognition evaluation, language understanding and speech synthesis, not only token counts.

Why Fertility Matters in Practice

High fertility has several concrete consequences.

  • Context window consumption. A model's context limit is measured in tokens. If Hindi text uses several times more tokens than English, the effective amount of conversation history, documents or instructions that fit is correspondingly smaller.
  • Latency per word. Autoregressive models generate one token at a time. If each Hindi word requires more tokens, producing the same spoken reply takes more generation steps, which matters in real-time settings such as voice conversations.
  • Throughput and compute. More tokens per request means more computation for both reading and generating, reducing how many requests a given system can serve.
  • Quality. Fragmented representations can make it harder for a model to learn word-level meaning and morphology, and sequence length limits during training mean fewer complete sentences of the language are seen per training example. The research cited above suggests tokenizer quality and downstream performance are related, though untangling cause from correlation with training data volume is difficult.

Mitigations in the Field

Researchers and practitioners have pursued several approaches.

Balanced tokenizer training

Training the tokenizer on a corpus that deliberately upsamples lower-resource languages gives them a fairer share of the vocabulary. The challenge is balancing this against efficiency for dominant languages, since vocabulary size is finite and larger vocabularies enlarge the embedding and output layers.

Vocabulary extension

An existing model's vocabulary can be extended with frequent Hindi, Gujarati or other Indic subwords. The new token embeddings must be initialised (a common heuristic is averaging the embeddings of the pieces each new token replaces) and the model then undergoes continued pretraining on Indic text so it learns to use them. Several open research efforts on Indian languages have reported substantial reductions in token counts with this approach, though the effect on quality depends on the amount and quality of continued training.

Indic-focused tokenizers and models

Models trained from the ground up with tokenizers designed for Indian languages can achieve low fertility across many scripts, often while respecting grapheme and conjunct boundaries. The Indian research community, including academic groups and public initiatives, has invested significantly in corpora and tooling for this.

Transliteration as preprocessing

Some systems transliterate between scripts, for example converting different Indic scripts into a common script to share vocabulary across related languages. This can improve sharing between Hindi and Gujarati, which are closely related, but it introduces conversion errors and requires reversing the transliteration faithfully on output.

Closing Thoughts

Tokenization is the first transformation any text undergoes in a language model, and its effects propagate through everything that follows. For Indian languages, the combination of corpus skew, three-byte UTF-8 encoding and the structure of abugida scripts means that the same meaning often costs more tokens than in English.

Practitioners building for Indian users should measure fertility on their own representative data, in the scripts their users actually write and speak, rather than assuming English-derived intuitions apply. Questions of script choice, normalization of spelling variants and model selection all benefit from that measurement. At Cirio, where conversations move fluidly between Hindi, Gujarati and English, this is one of many details we believe deserve careful attention, and the steady progress of Indic-focused research makes it an encouraging area to watch.

Frequently Asked Questions

What is token fertility?

Token fertility is the average number of tokens a tokenizer produces per word in a given language. A fertility close to one means most words map to a single token, while a high fertility means words are split into many fragments, which uses more of a model's context and compute.

Why does Hindi often produce more tokens than English?

Most widely used tokenizers learn their vocabulary from corpora dominated by English and other Latin-script languages, so fewer Devanagari sequences earn their own tokens. Devanagari characters also take three bytes each in UTF-8, so byte-level fallback multiplies the count further when a character sequence is unfamiliar.

Does writing Hindi in Roman script reduce token usage?

Romanized Hindi often tokenizes into fewer pieces because it reuses Latin-script subwords, but spelling is inconsistent across writers and the model may understand it less reliably. It is a trade-off between efficiency and consistency rather than a straightforward fix.

How can token fertility for Indic languages be improved?

Common approaches include training tokenizers on balanced multilingual corpora, extending an existing vocabulary with frequent Indic subwords and then continuing pretraining, and using tokenizers designed specifically for Indian languages. Each approach requires careful evaluation of downstream quality, not just token counts.

Put Voice AI to work for your business

Deploy an AI agent that handles calls in Hindi, English, and more in under a minute.

Start free: it's instant →