Back to Blog
Architecture6 min read1,158 words

Real Time Streaming Audio: How AI Voice Agents Speak Without Lag

The architecture behind sub 600ms AI responses, and every millisecond that matters.

Flowing teal and green waveform streams moving through dark space

The first time someone experiences a great AI voice agent, they are often surprised by how natural it feels. There is no awkward multi second pause after you speak. The agent picks up immediately, responds quickly, and sounds present in the conversation. This does not happen by accident: it is the result of a carefully engineered streaming audio pipeline where every component has been designed and tuned to minimize the time between the caller finishing a sentence and the AI beginning its response.

The simple approach to Voice AI treats the system as a loop: record the utterance completely, transcribe it, run it through an LLM, synthesize the response, and play it back. This works but creates terrible latency: a 5 second utterance plus 2 seconds of transcription plus 1 second of LLM response plus 2 seconds of TTS generation means the caller waits 10 seconds before hearing anything. Real production systems replace every batch operation with a streaming equivalent, achieving total response latencies under 600 ms.

The Pipeline: A Race Against Time

A Voice AI pipeline for phone conversations involves six major stages that run in sequence for each conversational turn. Each stage contributes latency that stacks on top of the PSTN and network delays we discussed in our telephony post. Understanding the pipeline is the first step to optimizing it, because you cannot improve what you cannot measure and model.

  • 1. Audio Capture: RTP audio received from the PSTN, chunked into 20 ms frames
  • 2. VAD: Voice Activity Detection identifies speech versus silence in real time
  • 3. ASR: Automatic Speech Recognition produces a rolling transcript as audio arrives
  • 4. LLM: Large Language Model generates response text given the full conversation
  • 5. TTS: Text to Speech synthesizes audio from the response text
  • 6. Playback: Synthesized audio sent back via RTP to the caller

Streaming ASR: Never Wait for Silence

Traditional batch ASR systems require the complete utterance before transcription begins. For a 10 second utterance, you wait 10 seconds before you even have the transcript to send to the LLM. Streaming ASR changes this fundamentally by producing rolling partial transcripts as audio arrives, frame by frame. The transcript is marked as interim (unstable) until end of turn is detected, at which point a final stable transcript is produced.

Leading streaming ASR systems like Deepgram Nova 2, AssemblyAI Streaming, and Whisper Live produce partial transcripts every 200 to 500 ms. This enables a powerful technique called speculative transcript feeding: rather than waiting for end of turn, a well designed Voice AI pipeline begins processing the transcript as soon as a complete meaningful phrase appears. If the caller says "I want to book an appointment for tomorrow" the LLM can begin reasoning about appointment booking before the caller finishes saying "around 3pm." This pipeline overlap is one of the biggest sources of latency reduction in modern Voice AI architectures.

Streaming LLM Inference

Modern LLMs support token streaming: rather than generating the complete response internally and then returning it all at once, they emit tokens one by one as they generate them. The time between sending the prompt and receiving the first token (called Time To First Token or TTFT) is the critical latency metric for Voice AI. For small fast models like Gemini Flash or GPT 4o mini, TTFT can be as low as 50 to 150 ms. For large frontier models under load, it can stretch to 500 to 800 ms.

The key insight for Voice AI is that you do not need the complete LLM response before starting TTS synthesis. You need only a natural speech boundary such as a period, a clause break, or a comma before you can synthesize and start playing the first chunk of audio. This means the AI can begin speaking its first sentence while the LLM is still generating sentences three and four. Implemented well, the caller cannot tell the difference between the AI knowing the full answer immediately versus assembling it on the fly.

Streaming TTS: The Last Latency Frontier

Traditional TTS systems require the complete text before synthesis begins; generating audio for a 30 word sentence might take 600 to 900 ms. Streaming TTS systems like ElevenLabs Streaming, Cartesia, and Deepgram Aura accept a token stream and begin synthesis immediately as words arrive. The first audio chunk from a streaming TTS system can appear in as little as 80 to 200 ms after receiving the first few words of text.

This means the AI agent can literally begin playing the first word of its response while still receiving text from the LLM for the remainder of the sentence. The overlap between LLM generation and TTS synthesis is where the most dramatic latency wins live. In a well optimized pipeline with a fast LLM and a low latency streaming TTS, the total TTFT for audio playback can reach 300 to 400 ms even accounting for PSTN network delay, which is well within the threshold where conversations feel natural.

Interruption Handling: The Hardest Part

Streaming the speech of the AI introduces a problem that batch systems do not face: what happens when the caller interrupts mid sentence? A production grade system must handle this gracefully and instantly. When VAD detects new speech from the caller while the AI is speaking, the system must simultaneously silence TTS playback, flush any queued audio chunks, cancel any in flight TTS generation requests, cancel any in progress LLM streaming if possible, and restart the ASR pipeline for the new utterance. This cascade of cancellation operations must happen in under 50 ms to feel seamless.

The engineering challenge is that all these components operate asynchronously. The TTS audio buffer might have 800 ms of audio queued up. The LLM might be mid sentence. Coordinating clean cancellation across all these asynchronous streams without leaking state and without the next turn starting with stale context is genuinely difficult. It is one of the areas where production Voice AI systems diverge most sharply from early demos.

Thinking in Latency Budgets

Every stage described above spends part of a single, shared budget: the time between a caller finishing a sentence and hearing the first syllable of a reply. Network transit, jitter buffering, speech detection, transcription, language model reasoning and speech synthesis all draw from it, and a gain in one stage is easily lost in another. The useful discipline is to measure each stage separately on real calls, look at the slow tail rather than the average, and treat the whole path as one system. As a rough guide to perception, replies that begin well under half a second feel nearly instant, while delays beyond a second or so start to make a conversation feel broken.

Frequently Asked Questions

How do AI voice agents achieve low latency responses?

AI voice agents achieve low latency by streaming every stage of the pipeline: streaming ASR produces partial transcripts as audio arrives, streaming LLM inference emits tokens as they are generated, and streaming TTS begins synthesizing audio before the full response is ready. Combined, this achieves total response latencies of 500 to 700 ms.

What is Time To First Token (TTFT) in voice AI?

Time To First Token (TTFT) is the time between sending a prompt to an LLM and receiving the first generated token. For voice AI, minimizing TTFT is critical because it directly contributes to how quickly the AI agent starts speaking after the caller finishes.

How do voice AI agents handle interruptions?

When a caller interrupts, the voice agent must simultaneously stop TTS playback, flush the audio queue, cancel in flight TTS and LLM generation, and restart ASR for the new utterance. This cascade must complete in under 50 ms to feel seamless to the caller.

What is speculative transcript feeding in voice AI?

Speculative transcript feeding is a technique where the LLM begins processing partial ASR transcripts before end of turn is detected. This overlaps ASR and LLM processing, reducing perceived response latency by 200 to 400 ms.

Put Voice AI to work for your business

Deploy an AI agent that handles calls in Hindi, English, and more in under a minute.

Start free: it's instant →