Library · Customer support and voice agents

Real-time AI: streaming, WebRTC and low-latency conversational agents

Engineer60 minUpdated: October 2026
64 of 105 in the library

Module: Voice & Real-Time AI | Time: ~25 min theory + 35 min practice


The gist

🎨 Picture this: Think about the difference between a letter and a phone call. With a letter, you write the whole thing, seal the envelope, take it to the post office and wait several days. The other person gets it all at once, in one piece. With a phone call, the words travel instantly; you hear the breathing, the pauses, the tone. Same information, completely different experience.

A batch request to an LLM is the letter: Claude thinks, assembles the entire answer, and only then sends it. Real-time streaming is the call: the first words show up on screen very quickly, while Claude is still "thinking" about the end of the sentence.

The difference in how it feels to the user is noticeable. In this lesson we'll go through how to build a streaming architecture, from a simple SSE endpoint all the way to a full voice agent with WebRTC.


Key concepts

  • Streaming: sending the LLM's answer word by word, as the tokens are generated, instead of waiting for the full response
  • Server-Sent Events (SSE): a one-way protocol. The server pushes data to the client over a regular HTTP connection, and the browser reads it with EventSource
  • WebSocket: a two-way protocol with a persistent connection. Both client and server can send data at any time; ideal for chats and interactive agents
  • WebRTC: the standard for peer-to-peer media streams (audio/video) in the browser with minimal delay; the foundation of modern voice agents
  • TTFT (Time To First Token): the time from sending a request to receiving the first token of the answer; a critical UX metric, and the lower the better
  • TPOT (Time Per Output Token): the average time to generate one token; affects how quickly text fills the screen
  • LiveKit: open-source WebRTC infrastructure for voice agents; people build voice products on it, and it has a Python framework called LiveKit Agents
  • OpenAI Realtime API: an API for two-way streaming audio; lets you build voice assistants without the intermediate STT → LLM → TTS steps. For the current realtime model and its prices, check OpenAI's documentation and pricing page (see also What's current)

Theory

Why streaming is critical for UX

A user staring at a blank screen for several seconds already thinks something broke. A user who sees the first words almost immediately and watches the text "type itself out" is engaged and perceives the system as alive.

Perceived speed matters as much as actual speed. Two services with the same total response time will feel different if one starts streaming immediately and the other holds a pause. For support chatbots and assistants, this usually shows up in user feedback; measure the effect on retention with your own data.

The Claude Streaming API: basic Python

The Anthropic SDK supports streaming out of the box. The key is using stream() instead of the regular messages.create():

python
import anthropic

client = anthropic.Anthropic()

# Streaming with a context manager
with client.messages.stream(
    model="claude-sonnet-5-5",   # current models: the What's current page
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain quantum entanglement in simple terms"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

# Get the final message after the stream
final_message = stream.get_final_message()
print(f"\n\nStop reason: {final_message.stop_reason}")
print(f"Input tokens: {final_message.usage.input_tokens}")
print(f"Output tokens: {final_message.usage.output_tokens}")

The text_stream method is the simplest option: it returns only the text deltas and ignores housekeeping events. If you need full control over events (content_block_start, ping, message_delta), use stream directly:

python
with client.messages.stream(...) as stream:
    for event in stream:
        if event.type == "content_block_delta":
            if event.delta.type == "text_delta":
                yield event.delta.text
        elif event.type == "message_stop":
            break

Server-Sent Events: streaming to the browser

SSE is the simplest way to get a streaming response to the browser. It's a regular HTTP GET that the server keeps open and periodically writes data into, in the format data: ...\n\n.

🎨 Picture this: SSE is like radio. You tune in to a station (open the connection) and listen. The broadcast is one-way: you can't talk back to the transmitter, only receive the signal.

FastAPI with the Anthropic SDK:

python
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
import anthropic
import json

app = FastAPI()
client = anthropic.Anthropic()

async def generate_stream(prompt: str):
    """Token generator for SSE"""
    with client.messages.stream(
        model="claude-sonnet-5-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": prompt}],
    ) as stream:
        for text in stream.text_stream:
            # SSE format: data: <payload>\n\n
            yield f"data: {json.dumps({'text': text})}\n\n"
    
    # End-of-stream signal
    yield "data: [DONE]\n\n"

@app.get("/stream")
async def stream_response(prompt: str):
    return StreamingResponse(
        generate_stream(prompt),
        media_type="text/event-stream",
        headers={
            "Cache-Control": "no-cache",
            "X-Accel-Buffering": "no",  # Turn off nginx buffering
        }
    )

The browser side uses the native EventSource:

javascript
const eventSource = new EventSource(`/stream?prompt=${encodeURIComponent(userInput)}`);
const output = document.getElementById('output');

eventSource.onmessage = (event) => {
    if (event.data === '[DONE]') {
        eventSource.close();
        return;
    }
    const { text } = JSON.parse(event.data);
    output.textContent += text;
};

eventSource.onerror = () => {
    eventSource.close();
    console.error('Stream ended or error occurred');
};

SSE limitations: GET requests only (long prompts won't fit), and no native support for two-way communication. For interactive chats you need WebSocket.

WebSocket: a two-way stream

WebSocket solves the SSE problem: the client can send messages at any time, and the server can reply with a stream. It's the foundation for chat interfaces.

python
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
import anthropic
import json

app = FastAPI()
client = anthropic.Anthropic()

@app.websocket("/ws/chat")
async def websocket_chat(websocket: WebSocket):
    await websocket.accept()
    conversation_history = []
    
    try:
        while True:
            # Get the user's message
            data = await websocket.receive_text()
            user_message = json.loads(data)
            
            conversation_history.append({
                "role": "user",
                "content": user_message["text"]
            })
            
            # Stream Claude's answer back over the WebSocket
            # (the sync client blocks the event loop; in production use AsyncAnthropic and async with)
            full_response = ""
            with client.messages.stream(
                model="claude-sonnet-5-5",
                max_tokens=2048,
                system="You are a helpful assistant. Answer in English.",
                messages=conversation_history,
            ) as stream:
                for text in stream.text_stream:
                    full_response += text
                    await websocket.send_json({
                        "type": "delta",
                        "text": text
                    })
            
            # End-of-answer signal
            await websocket.send_json({"type": "done"})
            
            # Add the answer to the history for the next turn
            conversation_history.append({
                "role": "assistant",
                "content": full_response
            })
    
    except WebSocketDisconnect:
        print("Client disconnected")

TTFT and TPOT: the metrics that matter

TTFT (Time To First Token) is the time from sending the request to the first token appearing. This is what the user feels as "the delay before the answer." Rough benchmarks (illustrative; tune them to your product):

  • Good: under half a second
  • Acceptable: about a second
  • Bad: noticeably more than a second

TPOT (Time Per Output Token) is the average time between tokens. It determines how "smoothly" text appears on screen. For comfortable reading, text should appear faster than a person reads it.

What affects TTFT:

  • Model size: as a rule, small models (Haiku) respond faster than large ones
  • Prompt length: long system prompts and history = slower
  • API load: peak hours = higher latency
  • Region: closer to the data center = less network latency
  • Prompt caching: cached tokens aren't recomputed, so for long contexts TTFT drops noticeably

Measure these metrics in production:

python
import time

start_time = time.time()
first_token_time = None
token_count = 0

with client.messages.stream(...) as stream:
    for text in stream.text_stream:
        if first_token_time is None:
            first_token_time = time.time()
            ttft = first_token_time - start_time
            print(f"TTFT: {ttft * 1000:.0f}ms")
        token_count += 1

total_time = time.time() - start_time
tpot = (total_time - first_token_time) / token_count * 1000
print(f"TPOT: {tpot:.1f}ms/token")

WebRTC and voice agents

🎨 Picture this: If SSE is radio and WebSocket is a phone line, then WebRTC is an intercom system inside a building. A direct connection, minimal middlemen, the voice travels almost instantly.

WebRTC (Web Real-Time Communication) is a browser standard for peer-to-peer audio and video. It's what Google Meet, Zoom on the web and Discord run on. For voice AI agents, WebRTC lets you:

  • Capture the user's microphone and stream the audio
  • Receive the agent's synthesized voice
  • Keep latency to a minimum (much lower than with regular HTTP requests)

The basic architecture of a voice agent:

Code
User speaks → [Microphone] → [WebRTC] → [STT: Whisper / Deepgram]
                                                        ↓
                                              [Claude (text)]
                                                        ↓
                                    [TTS: ElevenLabs / OpenAI TTS]
                                                        ↓
                        [User hears] ← [WebRTC] ← [Audio]

The problem with this architecture is that latency piles up at every step: speech recognition + Claude's first token + generation + voice synthesis. Together that can easily add up to several seconds, which is critical in a conversation.

LiveKit: infrastructure for voice

LiveKit is an open-source WebRTC server and SDK that handles the hard infrastructure parts: TURN/STUN servers, media routing, session recording. People build their own voice products on LiveKit; ready-made platforms like Vapi (see the lesson Voice AI agents) solve the same problem as a service.

A Python SDK voice agent with LiveKit:

python
from livekit import agents
from livekit.agents import AgentServer, AgentSession, Agent
from livekit.plugins import anthropic, openai, silero

class VoiceAssistant(Agent):
    def __init__(self):
        super().__init__(
            instructions="""You are a voice assistant. 
            Answer in English, briefly and to the point.
            Use images and analogies in your explanations."""
        )

server = AgentServer()

@server.rtc_session()
async def entrypoint(ctx: agents.JobContext):
    session = AgentSession(
        stt=openai.STT(language="en"),           # Speech-to-Text
        llm=anthropic.LLM(model="claude-sonnet-5-5"),  # LLM; current models: the What's current page
        tts=openai.TTS(voice="alloy"),           # Text-to-Speech
        vad=silero.VAD.load(),                   # Voice Activity Detection
    )
    
    # Noise cancellation and other room parameters are set through room_options
    # (see the LiveKit Agents documentation for your version)
    await session.start(room=ctx.room, agent=VoiceAssistant())
    
    # Greeting
    await session.generate_reply(
        instructions="Greet the user in English and offer to help"
    )

if __name__ == "__main__":
    agents.cli.run_app(server)

The plugins are installed separately (the livekit-agents, livekit-plugins-openai, livekit-plugins-silero and livekit-plugins-anthropic packages). The LiveKit Agents API has changed from version to version, so check the documentation.

LiveKit handles all the hard parts: echo cancellation, the jitter buffer, packet loss concealment. You focus on the agent's logic.

OpenAI Realtime API

OpenAI offers a fundamentally different approach: their Realtime API works over WebSocket (there's also a WebRTC option) and lets you send audio directly and get an audio answer back. There's no intermediate STT/TTS; the model handles voice natively. Event names and session fields changed when it came out of beta, so older examples from blogs may not work:

python
import asyncio
import json
import websockets
import base64

async def realtime_voice_session():
    # The model name in the URL is an example: take the current one from the model list in OpenAI's documentation
    url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1"
    headers = {
        "Authorization": f"Bearer {OPENAI_API_KEY}",
        # The OpenAI-Beta: realtime=v1 header isn't needed for the current version of the interface
    }
    
    async with websockets.connect(url, additional_headers=headers) as ws:
        # Configure the session. Audio format, voice and end-of-speech detection
        # are set in session.audio: see the Realtime API events reference for the exact fields
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "type": "realtime",
                "instructions": "You are a voice assistant. Keep your answers short.",
            }
        }))
        
        # Listen for events
        async for message in ws:
            event = json.loads(message)
            
            if event["type"] == "response.output_audio.delta":
                # Receive a chunk of the audio answer
                audio_data = base64.b64decode(event["delta"])
                # Play the audio...
            
            elif event["type"] == "response.done":
                print("Answer finished")

The advantage of the Realtime API is native voice processing without the STT → LLM → TTS chain, so the latency is lower than with a cascade and the conversation sounds more natural. The drawback: OpenAI models only. Claude isn't available in this format (use LiveKit with Claude for the cascade setup).

Reducing latency

A few techniques for minimizing latency in production:

1. Prompt caching. If your system prompt is long, cache it (the minimum length for caching depends on the model; see the documentation). Repeat requests won't recompute the prompt, and TTFT drops:

python
messages = client.messages.create(
    model="claude-sonnet-5-5",
    system=[{
        "type": "text",
        "text": long_system_prompt,
        "cache_control": {"type": "ephemeral"}  # Cache it
    }],
    messages=conversation,
    max_tokens=1024,
)

2. A smaller model for the first answer. The "cascade" strategy: start with Haiku (fast, cheap) and switch to Sonnet for complex questions.

3. Edge deployment. Cloudflare Workers AI or AWS Lambda@Edge: run inference closer to the user.

4. Streaming with early TTS. For voice agents, don't wait for the full text; pass each completed sentence to TTS:

python
buffer = ""
async for text in stream.text_stream:
    buffer += text
    # If we've collected a complete sentence, send it straight to TTS
    if any(buffer.endswith(p) for p in ['.', '!', '?', '...']):
        await tts_engine.synthesize_and_play(buffer.strip())
        buffer = ""

A real-world case: a streaming support chat

Picture this: an online store, hundreds of inquiries a day, customers waiting several minutes for an agent. You roll out a Claude agent with WebSocket streaming.

What changes: the user sees the first words of the answer almost immediately, and the wait stops feeling like a wait. The agent handles the routine questions, and human staff only see the complicated cases. What share of inquiries gets resolved without a human, and how much each answer costs, depends on your knowledge base and model: measure it in a pilot, and calculate the tokens using the prices on the What's current page.

The key architectural choice here is WebSocket instead of REST. One persistent connection per session, two-way exchange, no overhead of setting up a connection for every message.


Practice

Task: a streaming chat API with FastAPI and Claude

Goal: build a WebSocket server that streams Claude's answers in real time, with a simple HTML client.

Step 1: Install the dependencies

bash
pip install fastapi uvicorn anthropic websockets python-dotenv

Step 2: Backend, a FastAPI WebSocket server

Create a server.py file:

python
import os
import json
import asyncio
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
from fastapi.responses import HTMLResponse
import anthropic
from dotenv import load_dotenv

load_dotenv()

app = FastAPI(title="Streaming Chat API")
client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))

SYSTEM_PROMPT = """You are a helpful AI assistant. Answer in English, 
use structure and images in your explanations. Be brief."""

# A simple HTML client for testing
HTML = """
<!DOCTYPE html>
<html>
<head><title>Streaming Chat</title></head>
<body>
<h1>Streaming Chat with Claude</h1>
<div id="chat" style="height:400px;overflow-y:auto;border:1px solid #ccc;padding:10px;"></div>
<div style="margin-top:10px">
    <input id="input" type="text" style="width:80%" placeholder="Type a message..." />
    <button onclick="sendMessage()">Send</button>
</div>
<script>
    const ws = new WebSocket(`ws://${location.host}/ws/chat`);
    const chat = document.getElementById('chat');
    let currentBubble = null;
    
    ws.onmessage = function(event) {
        const data = JSON.parse(event.data);
        if (data.type === 'start') {
            currentBubble = document.createElement('div');
            currentBubble.style.cssText = 'margin:8px 0;padding:8px;background:#e8f4fd;border-radius:8px';
            currentBubble.innerHTML = '<strong>Claude:</strong> ';
            chat.appendChild(currentBubble);
        } else if (data.type === 'delta') {
            currentBubble.innerHTML += data.text;
            chat.scrollTop = chat.scrollHeight;
        } else if (data.type === 'done') {
            currentBubble = null;
        }
    };
    
    function sendMessage() {
        const input = document.getElementById('input');
        const text = input.value.trim();
        if (!text) return;
        
        const userBubble = document.createElement('div');
        userBubble.style.cssText = 'margin:8px 0;padding:8px;background:#f0f0f0;border-radius:8px;text-align:right';
        userBubble.innerHTML = `<strong>You:</strong> ${text}`;
        chat.appendChild(userBubble);
        
        ws.send(JSON.stringify({text}));
        input.value = '';
    }
    
    document.getElementById('input').addEventListener('keypress', (e) => {
        if (e.key === 'Enter') sendMessage();
    });
</script>
</body>
</html>
"""

@app.get("/")
async def get():
    return HTMLResponse(HTML)

@app.websocket("/ws/chat")
async def websocket_endpoint(websocket: WebSocket):
    await websocket.accept()
    history = []
    
    try:
        while True:
            data = await websocket.receive_text()
            user_msg = json.loads(data)
            
            history.append({
                "role": "user",
                "content": user_msg["text"]
            })
            
            # Start-of-answer signal
            await websocket.send_json({"type": "start"})
            
            full_response = ""
            
            # Stream Claude's answer
            with client.messages.stream(
                model="claude-haiku-4-5",  # Haiku for speed (check that the model is still available in the API)
                max_tokens=1024,
                system=SYSTEM_PROMPT,
                messages=history,
            ) as stream:
                for text in stream.text_stream:
                    full_response += text
                    await websocket.send_json({
                        "type": "delta",
                        "text": text
                    })
            
            # End of the answer
            await websocket.send_json({"type": "done"})
            
            # Save it to the history
            history.append({
                "role": "assistant",
                "content": full_response
            })
            
            # Limit the history (last 10 pairs)
            if len(history) > 20:
                history = history[-20:]
    
    except WebSocketDisconnect:
        print(f"Client disconnected, history length: {len(history)}")
    except Exception as e:
        print(f"Error: {e}")
        await websocket.close()

Step 3: Start the server

bash
# Make sure ANTHROPIC_API_KEY is in your .env file
uvicorn server:app --reload --port 8000

Open http://localhost:8000 in your browser. You'll see a chat interface with instant streaming.

Step 4: Measure TTFT in code

Add latency measurement to the WebSocket handler:

python
import time

start = time.time()
first_token = True
token_count = 0

with client.messages.stream(...) as stream:
    for text in stream.text_stream:
        if first_token:
            ttft = (time.time() - start) * 1000
            print(f"TTFT: {ttft:.0f}ms")
            first_token = False
        token_count += 1
        # ... send to the client

total = (time.time() - start) * 1000
tpot = total / token_count if token_count > 0 else 0
print(f"TPOT: {tpot:.1f}ms/token | Total tokens: {token_count}")

Step 5: Add an SSE endpoint for comparison

python
from fastapi.responses import StreamingResponse

@app.get("/stream")
async def stream_endpoint(prompt: str):
    async def generate():
        with client.messages.stream(
            model="claude-haiku-4-5",  # a fast model for real time
            max_tokens=512,
            messages=[{"role": "user", "content": prompt}],
        ) as stream:
            for text in stream.text_stream:
                yield f"data: {json.dumps({'text': text})}\n\n"
        yield "data: [DONE]\n\n"
    
    return StreamingResponse(
        generate(),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"}
    )

Test it: curl "http://localhost:8000/stream?prompt=What+is+AI%3F". You'll see the tokens arrive in real time right in your terminal.


Tools and resources

Tool Purpose Documentation
Anthropic SDK (Python) Streaming API with stream() and text_stream platform.claude.com/docs/en/build-with-claude/streaming
FastAPI WebSocket and SSE endpoints fastapi.tiangolo.com
LiveKit Open-source WebRTC for voice agents docs.livekit.io/agents
LiveKit Agents (Python) A framework on top of LiveKit with Claude integration github.com/livekit/agents
OpenAI Realtime API Native voice streaming developers.openai.com/api/docs/guides/realtime
websockets (Python) WebSocket client/server websockets.readthedocs.io
Deepgram Fast streaming speech recognition, an alternative to Whisper developers.deepgram.com
ElevenLabs High-quality TTS with streaming output elevenlabs.io/docs
uvicorn ASGI server for FastAPI www.uvicorn.org

Key takeaways

"Streaming isn't a technical detail, it's a UX decision. The user sees the first token almost immediately and feels the system is alive."

"TTFT is your main metric in real-time AI. If the first token arrives too late, the user loses patience. Cache your prompts and pick a fast model for the first touch."

"WebSocket for chat, SSE for one-way updates, WebRTC for voice: don't mix these patterns up. Each tool solves its own problem and has its own setup cost."


Next lesson

→ Chatbot managers: NLP, intent routing and smart escalation

The mark stays in this browser only and is never sent anywhere. My progress