Back to Blog
Language9 min read1,862 words

Code-Switching and Speech Recognition: The Challenge of Hinglish

Why the way most urban Indians actually speak is one of the hardest problems in speech recognition.

Abstract cover of two interleaved colour ribbons, one green and one amber, weaving through a soft sound waveform

Listen to a customer call anywhere in Mumbai, Ahmedabad or Delhi and you will rarely hear a sentence that belongs entirely to one language. A caller might say "mera order abhi tak deliver nahi hua, can you check the status?" and consider it a perfectly ordinary request. Linguists call this code-switching, and for speech recognition it is one of the most persistent unsolved problems in the Indian market.

The difficulty is not that Hindi is hard or that English is hard. Both have large bodies of research behind them. The difficulty is the combination: two sound systems, two vocabularies, two grammars and, crucially, several competing ways to write the result down. This post explains what code-switching is, why it defeats systems built for one language, and what general approaches the field has developed in response.

What Code-Switching Actually Is

Code-switching refers to a speaker alternating between languages within a conversation. Some researchers reserve the term code-mixing for switches inside a single sentence and use code-switching for switches across sentence boundaries, though in practice the terms are often used interchangeably. The useful distinction is structural:

  • Inter-sentential switching happens at sentence boundaries. "Main kal aaunga. Please keep the documents ready." Each sentence is internally monolingual.
  • Intra-sentential switching happens inside a sentence. "Aap mujhe ek callback de sakte ho kya, after 5 pm?" The languages meet within one clause.
  • Tag switching inserts a short tag or filler from another language: "Payment ho gaya, right?"

Intra-sentential switching is the hardest case for machines because the switch point can fall almost anywhere, including in the middle of a noun phrase or verb phrase.

The Matrix Language Frame Model

One influential account of how this mixing is organised is Carol Myers-Scotton's Matrix Language Frame (MLF) model. It proposes that in a mixed clause, one language acts as the matrix language, supplying the grammatical frame (word order and system morphemes such as inflections, postpositions and auxiliaries), while the other acts as the embedded language, contributing content words.

Hinglish illustrates this neatly. In "मेरा order cancel कर दो" (mera order cancel kar do), Hindi supplies the frame: the possessive mera, the light verb construction kar do, and verb-final word order. English contributes the content words order and cancel. The English verb does not take English inflection; it is slotted into a Hindi compound verb with karna. This pattern of English noun or verb plus a Hindi light verb (karna, hona, dena) is extraordinarily productive: "confirm kar dijiye", "block ho gaya", "refund de do".

The MLF model is not the only theory, and linguists debate its predictions, but it captures something practically important: code-mixed speech is not random. It has structure, and systems that model that structure do better than systems that treat it as noise.

How Common It Is in Indian Speech

Code-mixing between English and Indian languages is widespread in urban India, especially in commerce, technology, banking and customer service. Reliable large-scale measurements of exactly how often it occurs are hard to come by, and figures vary heavily with region, age, education, topic and setting, so we will not quote a number. What any practitioner can observe is that certain domains are saturated with English vocabulary regardless of the speaker's primary language: EMI, KYC, OTP, booking, delivery, account, balance, installation and warranty are routinely used by speakers who otherwise speak in Hindi or Gujarati.

This matters because customer service conversations sit squarely in those domains. A voice system serving Indian businesses that handles only pure Hindi or only Indian English will fail on a large share of real utterances.

Why Code-Switching Breaks Monolingual ASR

A conventional automatic speech recognition (ASR) system combines, explicitly or implicitly, three kinds of knowledge. Code-switching disrupts all three.

Acoustic Mismatch

The acoustic model maps audio to sound units. Hindi and English have different phoneme inventories. Hindi distinguishes aspirated and unaspirated stops (क vs ख, ka vs kha), dental and retroflex consonants (त vs ट), and breathy voiced stops (भ, घ). English has sounds such as the dental fricatives in think and this and a larger set of vowel contrasts. A model trained only on one language has no well-trained representation for many sounds in the other.

Lexical Gaps

A system's vocabulary, whether a word list or a subword inventory, reflects its training language. A Hindi-only system may have never seen cancel or refund as tokens, and will force them into the nearest Hindi-sounding words. An English-only system will do the same with kar do or nahi hua, often producing plausible but wrong English: "car though" for kar do is a classic failure.

Language Model Confusion

The language model, the component that predicts which word sequences are likely, is where code-switching hurts most. A monolingual model assigns very low probability to a Hindi postposition following an English noun, so even when the acoustics are clear, the decoder is pushed towards a monolingual interpretation. Modern end-to-end models blur the boundary between these components, but the underlying issue remains: they learn the statistics of their training data, and if mixed utterances are rare in that data, mixed utterances will be poorly recognised.

English Words with Indian Phonology

It is tempting to assume that the English words inside Hinglish sound like English words in an English corpus. They usually do not. Loanwords are adapted to the speaker's phonology. Station may be pronounced with a prothetic vowel as "istation". School may become "iskool". The /v/ and /w/ distinction is often merged. Word stress follows different patterns, and many speakers use retroflex consonants for English alveolar /t/ and /d/.

The result is that an Indian speaker's cancel is acoustically closer to other Indian speech than to American or British cancel. Systems trained on Indian English and on Indian-language speech that contains English loanwords handle this far better than systems that treat these words as foreign insertions.

The Script Problem: What Should the Transcript Look Like?

Even with perfect recognition, there is a question that monolingual ASR never has to answer: which script should the output use? Consider one utterance written three ways:

  • Devanagari throughout: "मेरा ऑर्डर कैंसल कर दो"
  • Roman throughout: "mera order cancel kar do"
  • Mixed script: "मेरा order cancel कर दो"

All three are defensible. Devanagari-only output is natural for Hindi readers but transliterates English words in ways that vary (ऑर्डर or आर्डर?). Roman-only output matches how many Indians type on phones, but romanised Hindi has no standard spelling: nahi, nahin and nhi all appear in the wild. Mixed-script output preserves each word's origin but requires the system to decide the language of borrowed words, and some words (time, phone, bus) are arguably part of everyday Hindi.

For Gujarati the same problem appears with a different script: "મારું order ક્યારે આવશે?" versus "maru order kyare aavshe?" Training data collected from different sources frequently mixes these conventions, which teaches models inconsistent output habits.

No Single Correct Transcript: The Evaluation Problem

The standard ASR metric, word error rate (WER), counts substitutions, insertions and deletions against a reference transcript. It assumes there is one correct reference. For code-mixed speech, that assumption fails. If the reference says "कैंसल" and the system outputs "cancel", a naive WER calculation counts an error even though the recognition is arguably perfect. Spelling variation in romanised Hindi creates the same false errors.

Researchers and practitioners respond in several ways:

  • Script normalisation before scoring, transliterating both hypothesis and reference into a common script so that script choice is not counted as an error.
  • Character-level or token-level metrics, such as character error rate, or mixed error rate variants used in other code-switching research, which reduce sensitivity to word segmentation.
  • Breakdown by language and switch point, since a system can look acceptable overall while failing almost entirely on the words immediately around a switch.
  • Task-level evaluation, asking whether the downstream intent or entity was captured correctly, which is often what a business actually cares about.

Beyond Hinglish: Gujarati-English and Hindi-Gujarati

Hinglish receives the most attention, but it is not the only mix. In Gujarat, Gujarati-English mixing follows similar patterns ("aa product nu warranty ketla time nu che?"). Many Gujarati speakers also move between Gujarati and Hindi, particularly in cities with large Hindi-speaking populations, and some conversations include all three languages. Hindi and Gujarati are closely related, share many cognates, and use scripts derived from a common ancestor, which makes language identification between them at the word level genuinely difficult. A word like paisa or kaam could plausibly belong to either.

Published research and public datasets for these combinations are considerably thinner than for Hindi-English, which in turn is thinner than for high-resource languages.

Data Scarcity

Good ASR needs large amounts of transcribed speech that matches the target condition. Code-mixed, telephone-quality, spontaneous Indian speech with consistent transcription conventions is scarce. Public efforts have helped: for example, the MUCS 2021 challenge at Interspeech included Hindi-English and Bengali-English code-switching tracks. But the volumes available are small compared with monolingual corpora, and read or scripted speech differs substantially from real customer conversations.

Collecting such data is also expensive. Transcribers must be bilingual, follow detailed script conventions and cope with the noise and disfluency of real calls.

General Approaches in the Field

No single technique solves code-switching, but several broad strategies are well established.

  • Multilingual models. Training one model on many languages, including Indian English and Indian languages, lets it share acoustic knowledge across languages and naturally represent mixed utterances. Large multilingual pretraining followed by fine-tuning on in-domain mixed speech is a common pattern.
  • Language identification. Some systems predict the language of each frame, word or segment, either as a separate component or as an auxiliary training objective, and use that signal to guide decoding.
  • Synthetic and augmented data. Researchers generate code-mixed text by applying linguistic constraints (including ideas inspired by the MLF model) to monolingual text, which helps train language models when real mixed text is scarce.
  • Transliteration and normalisation downstream. Before transcripts reach intent classification, entity extraction or search, they are often converted to a canonical form, for example a single script with standardised spellings, so that "cancel", "कैंसल" and "cancle" map to the same concept.
  • Domain adaptation. Biasing recognition towards known vocabulary such as product names, locality names and business terms reduces errors on exactly the words that matter most.

Looking Ahead

For Indian businesses, code-switching is not an edge case; it is the main case. The encouraging trend is that multilingual models are steadily improving, and evaluation practice is maturing towards script-aware and task-aware metrics.

For teams building or buying voice systems for Indian customers, the practical advice is straightforward: test on real, mixed, telephone-quality speech from your own domain, insist on knowing how transcripts are normalised before any accuracy figure is reported, and judge systems on whether they understood the caller, not merely on whether they spelled every word the way a particular reference did. At Cirio, we treat the way people actually speak as the starting point rather than an exception, and that is the standard any serious voice system in India should be held to.

Frequently Asked Questions

What is code-switching in speech?

Code-switching is the alternation between two or more languages within a single conversation, sentence or even phrase. In India it is extremely common in urban speech, where speakers routinely combine Hindi, Gujarati or other Indian languages with English words and phrases.

Why do speech recognition systems struggle with Hinglish?

Monolingual systems are trained on acoustic patterns, vocabularies and word sequences from one language, so switching mid-sentence produces sounds, words and word orders they rarely saw in training. Hinglish also has no single agreed writing system, which makes both training data and evaluation inconsistent.

Should a Hinglish transcript be written in Devanagari or Roman script?

There is no universally correct answer. Devanagari suits Hindi words, Roman script suits English words and matches how many people type, and mixed-script output preserves the origin of each word; the right choice depends on what downstream systems and readers need, as long as it is applied consistently.

How is code-switched speech recognition evaluated fairly?

Plain word error rate can penalise transcripts that are correct but written in a different script or spelling. Practitioners often normalise transcripts to a common script before scoring, report error rates separately for switch points and each language, and supplement automatic metrics with human review.

Put Voice AI to work for your business

Deploy an AI agent that handles calls in Hindi, English, and more in under a minute.

Start free: it's instant →