A language model learns to produce language by reading an enormous amount of it, and almost all of that language was written to be read. Web pages, documentation, articles and forum posts are built for a reader who can scan, skip back, glance at a heading and take in a bulleted list at a glance. So when you ask a model a question, its instinct is to produce a small document: a friendly opening, three bullet points, a bolded caveat and a summary line.
On a phone call, that instinct works against you. The caller cannot see bullets or bold text. They cannot re-read the sentence they missed. They hear a single stream of sound, once, in real time, often on a narrowband line with background noise. Designing prompts for a voice agent is therefore less about clever wording and more about moving the model from one register of language into another. This post covers what that shift involves and how to express it in instructions a model will actually follow.
Speech and Writing Are Different Registers
Linguists have studied the differences between spoken and written language for decades. Wallace Chafe described speech as tending toward "fragmentation" and "involvement", while writing tends toward "integration" and "detachment": writers pack more information into each clause using nominalizations, attributive adjectives and subordinate structures, because they have time to plan and the reader has time to decode. Douglas Biber's corpus work in Variation across Speech and Writing (1988) showed that the picture is more nuanced than a simple speech versus writing split, with many dimensions of variation, but face-to-face conversation consistently differs sharply from informational prose.
Chafe also observed that spoken language is delivered in short bursts, often called intonation units, that typically carry roughly one new piece of information each. That is a useful mental model for voice prompt design. A listener processes speech one chunk at a time. If a chunk carries three new ideas, something gets dropped.
The practical implications for a voice agent are straightforward:
- Short clauses, joined by simple connectors such as "and", "so" and "but"
- Concrete subjects and verbs rather than nominalized abstractions ("we can deliver it tomorrow" rather than "delivery is possible on the following day")
- Personal pronouns and direct address ("you", "your order")
- Minimal nesting: no clause inside a clause inside a parenthetical
Why Models Default to the Written Register
Beyond pretraining data, instruction tuning reinforces written habits. Assistants are typically rewarded for answers that look thorough and well organized on a screen, which means headings, numbered steps and comprehensive coverage. Chat interfaces render Markdown, so the model has learned that good answers contain asterisks and hashes.
When that output goes through a speech synthesizer, several things go wrong at once. Markdown symbols are either read aloud or silently stripped, and when they are stripped the structure they encoded disappears with them. A numbered list becomes a run-on sequence with no audible boundaries. Parenthetical asides lose their visual brackets and become confusing detours. A long sentence with its main verb at the end forces the listener to hold every preceding phrase in memory before the meaning resolves.
Information Structure: One Idea, Placed at the End
English grammarians describe a principle called end-focus: new or important information tends to come at the end of a clause, while given or known information comes first. Readers can tolerate violations because they can look back. Listeners cannot, so honouring end-focus matters more in speech.
Compare two ways of giving the same information:
- "Your appointment, which was originally scheduled for Tuesday but has been moved because the doctor is unavailable, is now on Thursday at 4 pm."
- "The doctor isn't available on Tuesday. So your appointment has moved to Thursday, at 4 pm."
The second version gives the reason first as context, then lands the new fact, the time, at the end of the final sentence where the listener's attention naturally rests.
Hindi and Gujarati are verb-final languages, so the main verb already arrives late, and it is easy for a generated sentence to accumulate long preverbal material before the listener knows what is happening. The same advice applies: keep the material before the verb short, and split sentences when the list of details grows. "Aapka order kal shaam tak pahunch jayega" (आपका ऑर्डर कल शाम तक पहुँच जाएगा) is easy to follow. A version with the order number, product name, warehouse and courier name all packed before "pahunch jayega" is not.
Numbers, Addresses and Email Addresses
Reading structured data aloud is where written-register output fails most visibly, and it is also where errors are most costly.
Numbers and amounts
A synthesizer given "₹12,499" may read it correctly, or it may produce something awkward depending on its text normalization. A prompt can reduce ambiguity by asking the model to write amounts the way a person would say them. Indian callers commonly use lakh and crore rather than hundred thousand and million, and a voice agent speaking to an Indian audience should follow suit: "one lakh twenty thousand rupees" rather than "one hundred and twenty thousand rupees".
Phone numbers and reference codes
Long digit strings should be grouped. A ten-digit mobile number is easier to follow in groups such as five and five, with a pause between groups. Reference codes that mix letters and digits benefit from spelling letters clearly, and from avoiding characters that sound alike on a phone line (B, D, P, and V are a classic confusion set). If you control the format of the codes, avoid those ambiguities at the source.
Email addresses and URLs
Email addresses should be spoken with symbols as words. In India, "at the rate" is a very common way to say "@", and a voice agent that uses it will sound natural to many callers. Periods become "dot", underscores become "underscore", and anything unusual should be spelled out.
Confirmation is part of the design
Whenever the agent captures structured data from the caller, it should read it back and ask for confirmation. This is not only politeness. Speech recognition errors on digits and names are common, and the read-back is the cheapest point at which to catch them. A good instruction is explicit about this: "After collecting a phone number, repeat it back in groups and ask the caller to confirm before continuing."
Brevity and Turn Length
In conversation, turns are short. People take the floor, make a point, and hand it back. A voice agent that delivers a six-sentence answer is not being thorough; it is holding the floor far longer than a human would, and callers respond by interrupting, disengaging, or asking the agent to repeat itself.
Useful principles for turn length:
- One main point per turn. If the caller asked two things, answer the first and offer the second, or answer both very briefly.
- End with a handover. A question or a clear prompt ("Shall I book that for you?") tells the caller it is their turn.
- Offer depth rather than forcing it. "There are three plans. Would you like me to go through them?" lets the caller decide whether they want the full list.
- Avoid filler preambles. Phrases like "That's a great question" add seconds of audio with no information.
Politeness and Register in Hindi and Gujarati
Register is not only about written versus spoken. It is also about social distance, and Indian languages encode that distance directly in grammar.
Hindi has three second-person pronouns: aap (आप, formal), tum (तुम, familiar) and tu (तू, intimate or, in the wrong context, rude). Verb agreement follows the pronoun, so the choice propagates through every sentence: "aap baithiye" (आप बैठिए), "tum baitho" (तुम बैठो). For business calls with customers, aap is the appropriate default. A model that drifts into tum halfway through a conversation, which can happen when the caller uses it, will sound disrespectful to many listeners.
Gujarati makes a similar distinction between tame (તમે), which serves as both the plural and the respectful form, and tu (તું) for familiar address.
Code-mixing is also a register choice. Many urban callers speak Hinglish naturally ("aapka payment successful ho gaya hai"), and a rigidly pure Hindi response full of formal Sanskrit-derived vocabulary can sound bureaucratic. On the other hand, some audiences expect more formal language. The prompt should state the intended register rather than leaving the model to guess.
Principles for Writing the Prompt Itself
Knowing what good spoken output looks like is half the work. The other half is instructing a model reliably. Some general practices that hold across models:
- Describe the medium, not just the style. "Your words will be converted to speech and heard over a phone call. The caller cannot see any text" gives the model a reason, which generalizes better than a list of banned characters.
- State positive instructions alongside prohibitions. "Do not use bullet points" is weaker than "When listing options, say how many there are, then name each one in a short sentence."
- Show a short contrasting example. A single pair of "written" versus "spoken" responses often communicates more than a paragraph of rules. Keep examples generic so the model does not copy their content verbatim.
- Specify formats for structured data. Tell the model exactly how amounts, dates, times and phone numbers should be written for the speech engine.
- Specify register and language behaviour. Which pronoun, whether to follow the caller's language switches, and how formal the vocabulary should be.
- Test by listening. Reading transcripts hides most problems. Play the synthesized audio over a telephone-quality channel and note every sentence you had to replay mentally.
A generic illustrative fragment might read:
You are speaking with a customer on a phone call. Everything you say
will be read aloud. Keep each reply to one to three short sentences.
Never use lists, symbols or formatting. When you give a number, write
it the way a person would say it. After you collect any detail from
the caller, repeat it back and ask them to confirm.Real deployments need instructions specific to the business, the audience and the tasks involved.
Designing for the Listener
The underlying shift is one of perspective. Written prompts implicitly optimize for a reader who controls the pace. Spoken output must optimize for a listener who does not. Every design decision follows from that: short units, important information at the end, structure carried in words rather than symbols, data read back in grouped form, and a register that fits the relationship between the business and its caller.
As voice agents take on more customer conversations in India, across Hindi, Gujarati, English and every mixture of them, this kind of careful language design becomes a core product skill rather than a finishing touch. At Cirio, we treat the question of how a response sounds as seriously as whether it is correct, because on a phone call those two things are never really separate. The best test remains the simplest one: close your eyes, listen to the response once, and ask whether you understood it the first time.
Frequently Asked Questions
Why do LLM responses sound unnatural when converted to speech?
Large language models are trained mostly on written text, so their default style uses long sentences, nested clauses, lists and visual formatting. These features help a reader scanning a page but overload a listener, who cannot re-read and must hold everything in working memory. Prompts for voice agents need to explicitly steer the model toward spoken register.
How should a voice agent read out phone numbers and email addresses?
It should break them into small, predictable groups, pause between groups, and spell ambiguous characters. For email addresses, it helps to say symbols as words, such as "at the rate" or "dot", which is how many Indian callers say them. Asking the caller to confirm afterwards catches transcription and hearing errors before they cause problems.
How long should a voice agent's response be?
Usually one to three short sentences, carrying one main point and ending with a clear handover such as a question. Long monologues are hard to follow on a phone line and make interruptions more likely. If more information is needed, it is better to deliver it across several turns and let the caller steer.
Should a Hindi voice agent use aap or tum?
For business calls, aap is the safe default because it signals respect toward a stranger or customer. Switching to tum can sound over-familiar or rude unless the brand and context clearly call for it. The prompt should state the expected register explicitly rather than leaving it to the model.