Listen carefully to any ordinary phone call and you will notice something that is easy to take for granted. One person stops, and the other starts, almost immediately. There is rarely a long silence and rarely a collision. The handover happens so smoothly that neither speaker is aware of performing a coordination task at all.
That smoothness is not an accident of politeness. It is one of the most tightly timed behaviours in human communication, and it is the benchmark every voice agent is implicitly judged against. A system that responds a second too late feels hesitant or broken; one that responds a moment too early feels rude. To understand why, we need to look at what decades of conversation research have found about how people actually take turns.
The Systematics of Turn-Taking
The modern study of turn-taking begins with a 1974 paper in the journal Language by Harvey Sacks, Emanuel Schegloff and Gail Jefferson, "A Simplest Systematics for the Organization of Turn-Taking for Conversation". Working from recordings of natural talk, they observed that conversation overwhelmingly follows a pattern of one speaker at a time, with transitions that minimise both gap and overlap.
Their model describes turns as built from turn-constructional units: sentences, clauses, phrases or even single words that can stand as a complete contribution in context. At the end of each such unit lies a transition-relevance place, a point where a change of speaker becomes possible. A small set of ordered rules then governs who speaks next: the current speaker may select the next speaker (for example, by asking a question to a named person), otherwise another participant may self-select, otherwise the current speaker may continue.
The key insight for engineers is that transition points are projectable. Listeners do not wait for silence to discover that a turn has ended. The grammar, meaning and prosody of the unfolding utterance let them anticipate where the end will fall.
The Gap Is Short Everywhere
For a long time, it was assumed that turn timing might vary enormously across cultures. Anecdotes suggested that some communities tolerate long silences while others talk over each other freely.
Stivers and colleagues tested this directly in a 2009 study in PNAS, "Universals and cultural variation in turn-taking in conversation". They examined responses to yes/no questions in video recordings from ten languages spanning very different cultures and language families. The result was striking: in every language, the distribution of response timing peaked near a short gap, with the most frequent offsets falling roughly between 0 and 200 milliseconds. Languages did differ in their average gaps (Japanese speakers were among the fastest and Danish speakers among the slowest in their sample), but these differences were on the order of a few hundred milliseconds, set against a shared underlying tendency to avoid both silence and overlap.
Other corpus studies, such as work by Heldner and Edlund (2010) on pauses, gaps and overlaps, point in the same direction while adding a nuance that matters a great deal for machines: speakers produce plenty of short overlaps and pauses, and pauses within a speaker's turn are often as long as, or longer than, the gaps between turns.
Planning Before the Turn Ends
The short gap creates a puzzle. Psycholinguistic research on speech production, summarised by Indefrey and Levelt (2004), indicates that going from a concept to the articulation of even a single word takes in the order of 600 milliseconds. Planning a full sentence takes longer. If listeners waited until the other person finished before starting to plan, typical gaps would be much longer than what we observe.
The resolution, argued in detail by Levinson and Torreira (2015) in Frontiers in Psychology, is that listeners do two things at once. While still comprehending the incoming turn, they predict its likely content and its likely end point, and they begin planning their own response in parallel. When the predicted end arrives and is confirmed by cues such as falling intonation or the completion of a syntactic unit, the prepared response is launched.
Experimental work supports this. De Ruiter, Mitterer and Enfield (2006) asked listeners to press a button when they anticipated the end of recorded turns. Removing pitch information had surprisingly little effect on accuracy, whereas making the words unintelligible degraded it substantially. The lexical and syntactic content of what is being said appears to carry much of the predictive load, with prosody contributing mainly near the very end.
Consider a caller who says: "Mujhe kal subah das baje ka appointment chahiye" (मुझे कल सुबह दस बजे का अपॉइंटमेंट चाहिए, "I need an appointment tomorrow at ten in the morning"). Hindi's verb-final word order is informative here. By the time "appointment" is heard, a human receptionist already knows roughly what is being asked and that a verb such as chahiye or karna hai is likely to close the turn. Planning of "Ji, das baje available hai" can begin before the caller finishes.
Why Silence Thresholds Are Reactive
Many voice systems decide that a user has finished speaking by waiting for a fixed stretch of silence after voice activity stops. This approach is simple and robust, but it is fundamentally reactive, and it fails in two predictable ways.
- It adds its own delay to every turn. A threshold can only fire after the silence has already elapsed. Whatever time the system waits is added on top of recognition, reasoning and speech synthesis, so the agent is guaranteed to start late relative to a human listener who was predicting the end.
- It confuses hesitation with completion. Callers pause while recalling an account number, deciding between options, or switching languages mid-sentence. A threshold short enough to feel responsive will cut people off during these pauses; one long enough to be safe will feel sluggish on every other turn.
This is a genuine trade-off rather than a tuning problem. No single silence duration resolves it, because the information needed to distinguish "thinking" from "done" is not in the silence. It is in what was said before the silence and how it was said.
Consider two utterances followed by the same pause:
- "Mera number hai nau aath do..." (मेरा नंबर है नौ आठ दो..., "My number is nine eight two...")
- "Haan, bas itna hi chahiye tha." (हाँ, बस इतना ही चाहिए था, "Yes, that was all I needed.")
A human knows immediately that the first speaker is mid-sequence and the second is finished. A silence detector treats them identically.
Gaps and Overlaps as Social Signals
Timing does not only coordinate who speaks. It carries meaning. Conversation analysts describe preference organisation: some responses, such as accepting an invitation or agreeing, are socially "preferred" and tend to come quickly, while declining, disagreeing or giving bad news tend to be delayed and prefaced. Kendrick and Torreira (2015) found in corpus data that longer gaps before responses are associated with dispreferred answers. Roberts, Francis and Morgan (2006) showed that listeners perceive longer silences before a response as signalling trouble or reluctance.
This has a direct consequence for machines. When a voice agent pauses for a long time before saying "Yes, that is available", it may unintentionally sound uncertain. When it pauses before a straightforward confirmation, a caller may start repeating themselves, which produces an overlap and a confusing exchange.
Overlap is also not uniformly bad. Much of it is cooperative: backchannels (a term introduced by Yngve in 1970) such as "hmm", "haan" (हाँ), "achha" (अच्छा) or "ji" (जी) in Hindi, and "haa" (હા) or "barabar" (બરાબર) in Gujarati, show attention without claiming the floor. Collaborative completions, where a listener finishes a phrase with the speaker, signal engagement. Treating every sound from the caller as an interruption misreads these signals entirely.
Cultural and Situational Variation
The cross-linguistic universals do not erase local differences. Deborah Tannen's work on conversational style described "high-involvement" speakers, for whom overlap signals enthusiasm, and "high-considerateness" speakers, for whom it signals rudeness. Norms also vary by setting. A customer calling a hospital enquiry line, a shopkeeper confirming a wholesale order, and an elderly caller unfamiliar with automated systems will each bring different expectations about pace and politeness.
In the Indian context, several factors add complexity:
- Code-mixing between Hindi, Gujarati and English often coincides with brief hesitations as speakers retrieve a word in another language.
- Honorific and politeness frames, such as addressing someone with ji or aap, can extend turns with elements that sound like closings but are not.
- Telephony conditions, including narrowband audio and network jitter, blur the fine prosodic cues that humans use near turn ends.
A design that works well for fast, transactional English speakers may feel abrupt to a caller who speaks more deliberately, and vice versa.
What This Implies for Voice Agents
The research does not prescribe a specific architecture, but it does point toward some general principles that the field has increasingly converged on.
Predict, do not merely detect
End-of-turn decisions benefit from combining acoustic evidence with linguistic evidence: whether the utterance so far forms a complete request, whether a question has been answered, whether a sequence such as a phone number is still in progress. Models that estimate the probability of turn completion from content and prosody behave far more like human listeners than fixed timers do.
Prepare responses early
Just as humans plan while listening, a system can begin interpreting and preparing a reply before it is certain the caller has finished, then commit or discard that work once the turn end is confirmed. The cost of occasionally discarding work is usually lower than the cost of always starting late.
Use acknowledgement tokens thoughtfully
A brief, natural acknowledgement such as "Ji" or "Achha, ek second" (अच्छा, एक सेकंड) can fill the space while a fuller answer is prepared, exactly as a human receptionist might. Used carelessly, though, such tokens become repetitive filler. They should sound appropriate to the language and register of the call.
Distinguish backchannels from barge-in
When a caller says "haan haan" while the agent is speaking, the agent should usually continue. When the caller says "nahi, ruko" (नहीं, रुको, "no, wait"), it should stop. Telling these apart requires attending to content, not just to the presence of sound.
Looking Ahead
The 200-millisecond gap is not a target to be hit by brute-force speed alone. It is the visible result of a predictive process: listeners model where a turn is heading, plan in parallel, and use timing itself to convey meaning. Voice agents that treat conversation as a sequence of isolated request and response pairs will always feel slightly out of step, however fast their individual components become.
The more promising direction is to build systems that listen the way people do, anticipating turn endings, reading hesitation and backchannels correctly, and adapting to the pace of the particular caller and language. This is the kind of problem we think about daily at Cirio, and it remains one of the most interesting open challenges in conversational AI. For businesses evaluating voice agents, a useful test is simple: call it, pause mid-sentence while reading out a number, say "haan" while it speaks, and see whether the conversation still feels like a conversation.
Frequently Asked Questions
How long is the typical gap between turns in human conversation?
Cross-linguistic research, most notably Stivers et al. (2009) in PNAS, found that the most common gap between a question and its answer is short, roughly in the range of 0 to 200 milliseconds. Languages differ in their averages, but the overall pattern of avoiding both long silences and overlap is remarkably consistent.
Why can humans respond faster than they can plan a sentence?
Psycholinguistic studies show that producing even a single word takes several hundred milliseconds of planning. People manage short gaps because they predict where the other speaker's turn will end and begin planning their reply while that turn is still in progress.
What is wrong with silence-based endpointing in voice agents?
A silence threshold can only fire after the speaker has already stopped, so it adds its full waiting time to every response. It also confuses mid-turn pauses, which are common when people think or recall details, with genuine turn endings.
What are backchannels and why do they matter for voice AI?
Backchannels are short listener signals such as 'hmm', 'haan' or 'achha' that show attention without claiming the turn. A voice agent must recognise them so it does not treat them as interruptions, and it can use similar tokens itself to signal that it has heard the caller.