Traditional TTS generates the entire audio before returning it. Streaming
returns audio chunks as they’re generated, providing:
- Lower latency: Playback can begin before the full utterance finishes. See Latency for measurement guidance.
- Better UX: Users hear audio immediately while more is being generated
- LLM integration: Process text token-by-token as it arrives from language models
The four rules
Streaming integrations live or die by these. Each links to the page that explains it in depth:- One turn per assistant reply, one session for the whole call. Keep the same streaming session open for the entire reply (never one session per sentence) and reuse it for the next reply. See Turn lifecycle.
- Send LLM tokens directly, without flushing. The server accumulates text
and starts generating at natural sentence boundaries. On
/ws/tts/streama flush ends the turn, so flushing mid-reply splits it into several turns, each with its own model prefill. See Chunking & per-segment latency. - Flush exactly once, at the end of the turn. This emits any trailing
text, then ends the turn.
endSession()/close()end the turn the same way. See Turn lifecycle. - Pre-connect at startup. Don’t pay the WebSocket handshake inside the first user interaction. See Latency.
Setup
Every snippet on this page uses aclient created once:
Simple streaming
The simplest pattern, streaming a complete text:- Python
- JavaScript
- Java
- cURL
LLM token streaming
Stream text token-by-token as it arrives from an LLM. Let the server handle chunking at sentence boundaries and flush once, when the reply ends.- Python
- JavaScript
- Java
- cURL
session comes from client.tts.streaming_session(...), opened once at
startup (see Complete agent turn). In Python,
send() returns only once the server has nothing more to deliver, so a
fast LLM waits on it while audio arrives. The complete example below reads
the LLM in a separate task so the two do not block each other.Complete agent turn
The full shape of a Python agent: one session opened at startup and reused for every reply, and the LLM read in its own task so a slowsend() never
stalls it.
Spelling out text mid-stream
Use<spell> tags to spell out text letter by letter. Spell content bypasses
text normalization automatically, while normalize: true still applies to
the surrounding text. Set an explicit language for character names:
</spell>
tag arrives before generating audio, and auto-closes incomplete tags if the
stream ends unexpectedly. See
Text processing for the full spell-tag
reference.
Audio playback
Error handling
Going deeper
Turn lifecycle
How turns start and end: flush, idle auto-flush, session reuse, usage
Chunking & per-segment latency
Flush granularity, buffer tuning, backpressure
Barge-in
Cancel the current turn when the user interrupts
Multi-context streaming
Up to 20 independent audio streams over one connection
Word timestamps
Word-level time alignments alongside streaming audio
WebSocket API reference
The full wire format: every message type, field by field