Skip to main content
Using a framework? Go to its guide: LiveKit, Pipecat, Vapi or Cognigy. The plugin handles the streaming loop for you. Building your own loop? Read on.
Traditional TTS generates the entire audio before returning it. Streaming returns audio chunks as they’re generated, providing:
  • Lower latency: Playback can begin before the full utterance finishes. See Latency for measurement guidance.
  • Better UX: Users hear audio immediately while more is being generated
  • LLM integration: Process text token-by-token as it arrives from language models

The four rules

Streaming integrations live or die by these. Each links to the page that explains it in depth:
  1. One turn per assistant reply, one session for the whole call. Keep the same streaming session open for the entire reply (never one session per sentence) and reuse it for the next reply. See Turn lifecycle.
  2. Send LLM tokens directly, without flushing. The server accumulates text and starts generating at natural sentence boundaries. On /ws/tts/stream a flush ends the turn, so flushing mid-reply splits it into several turns, each with its own model prefill. See Chunking & per-segment latency.
  3. Flush exactly once, at the end of the turn. This emits any trailing text, then ends the turn. endSession() / close() end the turn the same way. See Turn lifecycle.
  4. Pre-connect at startup. Don’t pay the WebSocket handshake inside the first user interaction. See Latency.

Setup

Every snippet on this page uses a client created once:

Simple streaming

The simplest pattern, streaming a complete text:

LLM token streaming

Stream text token-by-token as it arrives from an LLM. Let the server handle chunking at sentence boundaries and flush once, when the reply ends.
session comes from client.tts.streaming_session(...), opened once at startup (see Complete agent turn). In Python, send() returns only once the server has nothing more to deliver, so a fast LLM waits on it while audio arrives. The complete example below reads the LLM in a separate task so the two do not block each other.
Do not flush on every sentence from the client. On /ws/tts/stream every flush ends the turn, so send(token, flush=True) per sentence turns one reply into one turn per sentence, each paying a fresh model prefill. That makes latency worse, not better. See Chunking & per-segment latency.

Complete agent turn

The full shape of a Python agent: one session opened at startup and reused for every reply, and the LLM read in its own task so a slow send() never stalls it.

Spelling out text mid-stream

Use <spell> tags to spell out text letter by letter. Spell content bypasses text normalization automatically, while normalize: true still applies to the surrounding text. Set an explicit language for character names:
When streaming token-by-token, spell tags that span multiple chunks are handled automatically: the server buffers text until the closing </spell> tag arrives before generating audio, and auto-closes incomplete tags if the stream ends unexpectedly. See Text processing for the full spell-tag reference.

Audio playback

Error handling

Going deeper

Turn lifecycle

How turns start and end: flush, idle auto-flush, session reuse, usage

Chunking & per-segment latency

Flush granularity, buffer tuning, backpressure

Barge-in

Cancel the current turn when the user interrupts

Multi-context streaming

Up to 20 independent audio streams over one connection

Word timestamps

Word-level time alignments alongside streaming audio

WebSocket API reference

The full wire format: every message type, field by field
Last modified on September 22, 2026