Retrieval-augmented generation, usually shortened to RAG, has become the default answer to a simple question: how do you make a general-purpose language model answer accurately about one specific business? Instead of hoping the model knows your return policy, you find the relevant passage at query time and put it in front of the model before it answers.
In text applications, a retrieval step that adds a noticeable pause is often acceptable. A chat interface can show a typing indicator. On a live phone call, there is no typing indicator. Every moment spent retrieving is heard by the caller as silence, and silence in conversation carries meaning: hesitation, confusion or a dropped line. This post looks at how real-time constraints reshape the design of RAG, conceptually and without prescribing any particular configuration.
RAG in Brief
The idea was formalised in the 2020 paper by Lewis and colleagues that introduced the term, though the broader pattern of combining search with generation predates it. A typical RAG system has three stages:
- Indexing. Source documents such as FAQs, policies, product catalogues and manuals are split into passages (chunks) and indexed for search.
- Retrieval. When a question arrives, the system searches the index and returns the passages most likely to contain the answer.
- Generation. The retrieved passages are placed in the model's context together with the conversation, and the model composes a response grounded in them.
The appeal is practical. Knowledge can be updated by editing documents instead of retraining a model, answers can be traced back to sources, and the model is less tempted to invent business-specific facts.
Why a Live Call Changes the Design
Human conversation has a rhythm. Turn transitions between speakers are typically short, and listeners notice when a reply is delayed. A voice agent's total response time is the sum of many steps: detecting that the caller has finished, transcribing the speech, deciding what to do, retrieving knowledge, generating a reply and synthesising audio. Retrieval is one link in that chain, and the whole chain has to feel conversational.
This has several design consequences:
- Retrieval must be predictable, not only fast on average. An occasional slow search is heard as an awkward pause. Tail behaviour matters as much as typical behaviour.
- Retrieval needs a fallback. A system should know what to do if retrieval does not return in time: answer from the conversation so far, acknowledge and continue, or defer the question gracefully.
- Every stage competes for the same time. Richer retrieval (more candidates, extra re-ranking, query rewriting with a model) improves quality but consumes time that could otherwise go to generation or speech.
- Queries are spoken, not typed. They arrive through speech recognition, may contain recognition errors, and are phrased conversationally, often depending on earlier turns: "aur uska warranty kitna hai?" (and what is its warranty?) means nothing without knowing what "uska" refers to.
The right balance depends on the domain, the size of the knowledge base and the infrastructure. What is constant is that retrieval design for voice is an exercise in trade-offs rather than maximising any single metric.
Lexical, Dense and Hybrid Retrieval
Lexical Retrieval
Lexical methods match the words in the query against the words in documents. The best-known scoring function is BM25, developed from the probabilistic relevance framework associated with Stephen Robertson and colleagues. BM25 rewards documents that contain query terms, gives more weight to rare terms than common ones, and normalises for document length.
Lexical retrieval is fast, well understood and excellent at exact matches: product codes, model numbers, branch names, scheme names. Its weakness is vocabulary mismatch. A caller who asks "paise wapas kab milenge?" (when will I get my money back?) may not share a single word with a document titled "Refund Timelines".
Dense Retrieval
Dense methods encode queries and passages as vectors (embeddings) using a neural model trained so that texts with similar meaning land close together. Retrieval becomes a nearest-neighbour search in vector space, typically accelerated with approximate nearest-neighbour indexes.
Dense retrieval handles paraphrase and synonymy well, and multilingual embedding models can match a query in one language with a passage in another. Its weaknesses are the mirror image of lexical search: it can be imprecise with exact identifiers, numbers and rare proper nouns, and its behaviour is harder to inspect.
Hybrid Retrieval and Reciprocal Rank Fusion
Because the two families fail in different ways, many production systems combine them. The practical challenge is merging two ranked lists whose scores are on incomparable scales. A widely used, simple answer is reciprocal rank fusion (RRF), described by Cormack, Clarke and Buettcher in 2009. RRF ignores raw scores and uses only positions: each document receives a contribution based on the reciprocal of its rank in each list (with a smoothing constant), and the contributions are summed.
RRF is attractive in real-time settings because it is cheap, needs no training, and is robust to differences in score distributions. More elaborate fusion or learned re-ranking can improve quality further, at the cost of additional computation.
Multilingual Retrieval in India
Indian deployments add a layer of complexity that most RAG literature does not address directly.
- Query and document languages differ. A caller asks in Hindi or Gujarati, but the knowledge base is written in English, as business documentation in India very often is.
- Script varies. The same query may be transcribed in Devanagari ("रिफंड कब मिलेगा"), in Roman script ("refund kab milega") or in a mixture, depending on the speech recognition system.
- Code-mixing is normal. Queries routinely mix English domain terms with Indian-language grammar: "EMI option available hai kya is phone pe?"
- Transliterated names are inconsistent. Locality and product names can be spelled many ways in Roman script.
General strategies include multilingual embedding models that map different languages into a shared space, normalising transcripts to a consistent script before search, maintaining key content in the languages callers actually use, and relying on lexical matching for English domain terms that survive intact inside mixed-language queries. Hybrid retrieval is particularly helpful here, since the lexical side catches "EMI" while the dense side captures the meaning of the Hindi around it.
Chunking for Answers That Will Be Spoken
How documents are divided into passages affects both retrieval quality and the answer that ultimately reaches the caller's ear.
- Self-contained chunks. A passage should make sense on its own. A chunk that says "this applies only to the above plans" is useless if the plans are in a different chunk.
- One idea per chunk, where possible. Focused passages retrieve more precisely and give the model less irrelevant material to wade through.
- Keep structure meaningful. Tables and bullet lists that look fine on a screen can become confusing when flattened into text. Converting them into clear statements often helps.
- Write for the ear. Spoken answers should be short. A knowledge base written as dense paragraphs of conditions and exceptions invites long, hard-to-follow spoken replies. Source content that leads with the direct answer tends to produce better voice responses.
When Not to Retrieve
A great deal of any phone conversation does not need knowledge retrieval at all:
- Greetings and small talk: "Namaste", "haan boliye", "thank you".
- Confirmations and acknowledgements: "haan sahi hai", "okay", "ek minute".
- Collecting information the caller is providing, such as a name or preferred time.
- Follow-ups that can be answered from what was already retrieved earlier in the call.
Retrieving on every turn wastes time and can actively harm quality, since irrelevant passages in context may distract the model. Deciding when retrieval is warranted, and reusing relevant context across turns, is one of the most effective ways to make RAG compatible with conversational timing.
Stale Knowledge
RAG is only as accurate as its sources. A beautifully tuned retrieval system will confidently serve last season's prices if nobody updated the document. Common problems include duplicated documents with conflicting versions, outdated offers that were never removed, and information that changes daily (stock, schedules) being stored as static text.
Good practice includes clear ownership of each knowledge source, removing superseded content rather than adding new content alongside it, attaching validity dates to time-sensitive information, and fetching genuinely dynamic data from live systems instead of a document index.
Evaluating Retrieval and Answers
It is useful to evaluate retrieval and generation separately, because a wrong answer can come from either.
Retrieval Metrics
- Recall at k asks whether the passage containing the answer appears among the top k results. If it does not, the model never had a chance.
- Mean reciprocal rank and similar ranking metrics reward placing the right passage near the top.
- Test sets should reflect real callers, including spoken phrasing, recognition errors, Hindi and Gujarati queries and code-mixed questions, not only clean English questions written by the team that wrote the documents.
Answer Metrics
- Faithfulness checks whether every claim in the answer is supported by the retrieved passages.
- Answer relevance checks whether the response actually addresses the question.
- Appropriate abstention checks whether the agent declines when the knowledge base does not contain an answer.
Automated judging with language models can scale these checks, but it should be calibrated against human review, especially for Indian languages where automated judges may be less reliable.
Looking Forward
RAG remains one of the most practical ways to make voice agents knowledgeable about a specific business. Real-time conversation does not change the fundamentals, but it changes the priorities: predictability over peak quality, knowing when not to retrieve, handling multilingual and spoken queries, and content that works when read aloud.
For teams building voice systems, the best starting point is an honest evaluation set drawn from real calls. It reveals quickly whether problems lie in retrieval, in the knowledge itself or in generation, and it keeps design decisions grounded in how callers actually ask questions.
Frequently Asked Questions
What is retrieval-augmented generation?
Retrieval-augmented generation (RAG) is a technique in which relevant passages are retrieved from a knowledge source and supplied to a language model alongside the user's question. The model then generates an answer grounded in that retrieved text rather than relying only on what it learned during training.
Why is RAG harder in a voice agent than in a chatbot?
On a phone call, any delay before the agent responds is heard as silence, and callers find long pauses unnatural. Retrieval therefore has to fit inside a tight conversational rhythm, which affects how queries are formed, which retrieval methods are used and when retrieval is skipped altogether.
What is hybrid retrieval?
Hybrid retrieval combines lexical search, which matches exact words and is good for names, codes and numbers, with dense search, which matches meaning through embeddings. The two ranked lists are merged, for example with reciprocal rank fusion, so each method covers the other's blind spots.
How do you evaluate a RAG system?
Evaluation usually separates retrieval quality from generation quality. Retrieval is measured with metrics such as recall at k, which checks whether the right passage appears in the top results, while generation is checked for faithfulness, meaning whether the answer is actually supported by the retrieved passages.