Back to Blog
Conversation Science9 min read1,825 words

Backchannels and Barge-In: Managing Interruptions in Voice Conversations

Why a voice agent must tell the difference between a listener saying haan and a caller trying to take the floor.

Abstract illustration of two overlapping sound waves in teal and white, one small and rhythmic, one rising to cut across the other

Listen closely to any phone call between two people in India and you will hear a steady stream of small sounds from whoever is not speaking: haan, ji, hmm, achha, theek hai. None of them is an attempt to take over the conversation. They are a way of saying "I am here, keep going." Remove them and the speaker soon asks, "Hello? Aap sun rahe ho?"

For a voice agent, these sounds create a hard problem. The agent is speaking, and suddenly there is caller audio on the line. Is that a listener nodding along, or a caller who wants to correct the address the agent just read back? Stop too eagerly and the agent becomes a nervous speaker who abandons every sentence. Stop too reluctantly and it talks over a customer who is trying to say something important. This post covers what linguistics tells us about backchannels and overlap, how barge-in evolved in telephony systems, and the general principles for handling interruptions well.

What linguists mean by a backchannel

The term back channel was coined by Victor Yngve in his 1970 paper "On getting a word in edgewise." His observation was that a conversation has two channels running at once: the main channel carried by the person holding the turn, and a back channel through which the listener sends short feedback without claiming the turn.

Later conversation analysis refined this. Emanuel Schegloff described certain tokens as continuers: signals like mm hm that explicitly pass up the opportunity to speak and invite the current speaker to continue. Researchers have since distinguished several functional types, and while the boundaries are debated, a useful working taxonomy is:

  • Continuers: minimal tokens that say "go on." Hmm, haan, ji.
  • Acknowledgements: tokens that register receipt of information. Achha, theek hai, okay, Gujarati haa or saru.
  • Assessments: tokens that evaluate what was said. Bahut badhiya, arre wah, sahi hai, Gujarati barabar (right, exactly).

The distinction matters because the function shifts along this scale. A continuer almost never wants the floor. An acknowledgement may be closing a topic. An assessment can be the first move of a new turn: "Sahi hai, lekin mujhe Saturday chahiye" (fine, but I need Saturday) begins as agreement and turns into a correction.

Backchannels in Indian conversation

Backchannel frequency and form vary by culture and language, and Indian phone conversations tend to be rich in them. Some patterns a voice designer should expect:

  • Honorific continuers. Ji and haan ji (हाँ जी) are extremely common, especially when speaking to a business or someone perceived as senior. Their politeness can mislead a system: ji sounds like assent, but it often just means "I am listening."
  • Gujarati tokens. Haa (હા), barabar (બરાબર), saru (સારું) and hmm appear where a Hindi speaker would say haan or theek hai.
  • Code-mixed feedback. Okay okay, right, haan okay, achha okay. English tokens mixed with Hindi are normal in urban speech.
  • Non-lexical sounds. Breath-like hm, tongue clicks and short laughs carry feedback too, and they are the hardest for a recogniser to classify.

Cooperative and competitive overlap

Not all overlapping speech is a problem. Conversation analysts distinguish cooperative overlap, where the listener's speech supports the current speaker, from competitive overlap, where the incoming speaker is trying to take the turn. Work by Peter French and John Local in the early 1980s on "turn-competitive incomings" showed that competitive interruptions in English tend to be marked prosodically: they are typically louder and higher in pitch than the surrounding speech, as if the speaker is raising their voice to win the floor.

Cooperative overlap includes backchannels, but also collaborative completions (finishing the other person's sentence) and choral agreement. Competitive overlap includes corrections ("Nahi nahi, Andheri West"), urgent questions, and attempts to end the call.

Humans resolve overlap quickly. Usually one speaker drops out within a syllable or two, and the remaining speaker sometimes recycles the words that were buried in the overlap. A voice agent must make the same decision, but with far less information and a noisier signal.

Barge-in: history and failure modes

A short history

Early interactive voice response systems played prompts to completion and only then listened. Callers who already knew the menu had to wait through "For account balance, press 1" every time. Barge-in (sometimes called cut-through or speak-through) was introduced so that a DTMF keypress, and later a spoken command, would stop the prompt immediately.

Making this work on a phone line required solving a technical prerequisite: echo. On many telephone connections, part of the system's own outgoing audio returns on the incoming path, through line echo at hybrid circuits or acoustic echo from a speakerphone. Without echo cancellation, the recogniser hears the prompt itself and "barges in" on its own voice. Echo cancellation became a standard component of barge-in capable telephony platforms for this reason.

Classic IVR barge-in was also relatively simple because the expected input was narrow: a digit, "yes," "agent," a menu option. Grammar-based recognisers could require that audio match an expected phrase before stopping the prompt. Open-ended voice agents do not have that luxury. The caller can say anything, so the question "is this a real interruption?" cannot be answered by matching against a short list.

Why false barge-in happens

False barge-in is when the system stops speaking even though the caller did not intend to take the turn. Common causes include:

  1. Backchannels. The most frequent case in natural conversation. The caller says haan ji to be polite and the agent abandons its sentence.
  2. Residual echo. Imperfect cancellation, especially with speakerphones, mobile handsets in cars, or network paths that add delay, lets fragments of the agent's own voice leak back.
  3. Background speech. A television, a shopkeeper talking to a customer, family members at home, a public announcement at a railway station. This audio is real speech, so a detector tuned only to "is there speech?" will fire.
  4. Non-speech noise. Traffic, pressure horns, a fan close to the microphone, packet loss artefacts and comfort noise on narrowband lines.

The opposite failure, missed barge-in, is just as damaging. A caller says "Nahi, galat number hai" and the agent continues reading an entire paragraph. Callers experience this as not being listened to, and they often escalate their voice or hang up.

The design dilemma: too sensitive or too deaf

Every interruption policy sits somewhere on a trade-off curve. Make the system react to the faintest caller audio and it will stop for every hmm and every passing autorickshaw. Make it require strong evidence and it will steamroll callers who have something real to say.

There is no single correct point on this curve, because the cost of each error depends on context:

  • While the agent is reading back critical details (an amount, an address, an OTP instruction), missing a correction is costly, so leaning towards sensitivity makes sense.
  • While the agent is delivering a long explanation in a quiet moment, spurious stops are more costly than a slight delay in yielding.
  • Right after the agent asks a question, almost any caller speech is likely to be an answer, and the agent should generally be close to finishing anyway.

This is why well-designed systems treat interruption handling as a decision informed by conversational context, not as a fixed audio switch.

What signals a system can use

Systems in research and industry draw on several families of evidence. None is sufficient alone.

Duration and persistence

Backchannels are short. A real interruption usually continues. Waiting to see whether caller speech persists is the simplest discriminator, but every moment of waiting is a moment the agent keeps talking over the caller, so it trades accuracy against responsiveness.

Lexical content

Once some words are recognised, their content is informative. A lone haan, hmm, ji or okay is likely a backchannel. Words like nahi, ruko, wait, lekin, sorry, or any content-bearing phrase ("mera order number...") suggest a real turn. The challenge is that recognition of very short, overlapped, code-mixed utterances is itself error-prone, and waiting for words adds delay.

Prosody

Pitch, energy and speaking rate carry turn-taking information. Following the work on competitive incomings, a louder, higher-pitched onset is more likely to be a bid for the floor. Research by Nigel Ward and Wataru Tsukahara also studied the reverse direction: prosodic cues in the speaker's voice, such as a region of low pitch, that invite backchannels from a listener. Prosodic cues vary across languages and speakers, so they work best as one input among several.

Dialogue state

What the agent was doing matters. Was it mid-sentence or near the end? Did it just ask a yes or no question? Is it reading out a value the caller might want to correct? The same haan means something different in each situation.

Recovering gracefully after an interruption

Deciding to stop is only half the problem. What the agent does next determines whether the interruption feels natural.

Track what was actually heard. If the agent had planned to say "Aapka appointment Tuesday ko 4 baje hai, aur doctor ka naam Dr. Mehta hai" and was cut off after "Tuesday ko," the conversation history should not record the whole sentence as delivered. Otherwise the agent will later behave as if the caller knows the time and the doctor's name. Keeping the delivered portion separate from the generated portion is a basic requirement for coherent recovery.

Respond to the interruption first. If the caller said "Tuesday nahi, Wednesday," the next turn should address that directly rather than finishing the old sentence.

Repeat only what still matters. After handling the interruption, the agent can re-offer the unheard information if it is still relevant, preferably rephrased rather than replayed word for word.

Handle false stops smoothly. If the agent paused and the caller audio turns out to have been a backchannel, the natural move is to continue, much as a human speaker picks up after a listener's achha. Restarting the whole sentence from the beginning sounds robotic.

Looking ahead

Turn-taking research is moving from rigid, sequential models towards systems that listen continuously while speaking, closer to how people actually converse. Full-duplex dialogue models, richer prosodic modelling and training data drawn from real multilingual phone calls all point in the same direction: agents that understand that haan ji is a nod, not a demand.

For teams building voice agents for Indian callers today, the practical lesson is to treat interruptions as a conversational phenomenon rather than an audio threshold. Collect real calls, label backchannels and genuine interruptions separately, measure both false and missed barge-in, and review how the agent recovers afterwards. At Cirio, this is the kind of detail we believe separates an agent that callers tolerate from one they are comfortable talking to.

Frequently Asked Questions

What is a backchannel in conversation?

A backchannel is a short signal a listener produces while another person holds the floor, such as hmm, haan, ji or achha. It shows attention or agreement without trying to take over the turn. The term was introduced by the linguist Victor Yngve in 1970.

What is barge-in in a voice system?

Barge-in is the ability of a caller to interrupt a system while it is speaking, causing the system to stop its prompt and listen. It was introduced in IVR systems so experienced callers would not have to wait through long menus. Modern voice agents need the same ability, but with more judgement about what counts as a real interruption.

Why do voice agents stop talking for no reason?

This is usually false barge-in. The system mistakes something that is not a real interruption, such as a backchannel, background television, line noise or an echo of its own voice, for a caller trying to speak. The result is a halting conversation where the agent keeps breaking off mid-sentence.

How should a voice agent respond after being interrupted?

It should track how much of its own response was actually heard before it stopped, rather than assuming the whole sentence was delivered. It should then address what the caller said, and only repeat or rephrase the unheard part if it is still relevant.

Put Voice AI to work for your business

Deploy an AI agent that handles calls in Hindi, English, and more in under a minute.

Start free: it's instant →