Skip to main content
/ws/tts/stream is one logical TTS request per turn, regardless of how many send calls you make. The server’s text buffer accumulates tokens and hands each sentence to the model as soon as its closing ., ? or ! arrives. Text with no sentence end yet is synthesized after 2 s without new input, so end each sentence with punctuation or flush at the end of the reply. Inside a single turn, model state (KV cache, voice conditioning) is preserved across chunks so prosody stays natural. On /ws/tts/stream, a flush ends the turn. The server synthesizes what is buffered, sends final and then session_closed with its own usage block, and the next text starts a new turn on a fresh engine session. There is no mid-turn flush: flushing per sentence turns one reply into N turns, each with its own model prefill (the full model time-to-first-audio, see Latency), its own final/session_closed and its own usage record. See Turn lifecycle.

Flush granularity: flush once per reply

What decides time-to-first-audio per segment is how often you flush, not how you split the text you send. Raw LLM tokens sent without flush are fine: the server’s buffer reassembles them and only hands sentence-sized work to the model. If the full text is available before TTS starts, send it in one message with flush. We deliberately don’t publish exact ms figures here: they depend on region, voice, and deployment. To reproduce the comparison for your own deployment, time each chunking strategy against your endpoint as described in Measuring TTFA correctly.

Server-side chunking

Native streaming sessions use the server’s sentence-aware text buffer. Send tokens as they arrive and let that buffer decide when a complete unit is ready. Two message fields tune it (see the Stream Input reference):
  • flush_timeout_ms (default 500): the longest a period waits when only the next token can tell what it means: after a single letter (z. could become z.B.) or after a number (1299. could become 1299.99, 3. could become 3. Oktober); after a number it waits up to twice as long. The next token ends the wait. Every other sentence is synthesized as soon as it ends. Text with no sentence end yet waits 2 s.
  • max_buffer_length: the most characters the buffer holds before it is synthesized regardless of boundaries.
The native API has no diffusion-step or client-defined chunk-schedule controls. ElevenLabs-style fields such as chunk_length_schedule or auto_mode are unknown to it: the server ignores them and replies with a warning frame. The SDK code for this pattern is in the Streaming overview.

Handle backpressure

If audio arrives faster than you can play it, bound your buffer instead of letting it grow:

Common mistakes

  • Per-sentence flush=true. Every flush ends the turn, so the next text pays a fresh prefill and produces its own session_closed and usage record. Flush once per reply.
  • One session per sentence. A new WebSocket handshake plus a fresh model prefill, every sentence. Keep the same session open for the whole assistant reply and reuse it for the next one. See Turn lifecycle.
  • Client-side sentence buffering before send. Unnecessary: the server already buffers tokens and chunks at sentence boundaries. Pre-buffering on the client just adds latency.
  • Calling send(text, flush=true) per word “for lower latency.” It is the opposite: each flush is a separate turn and model call. Word-granular flushing produces the worst possible TTFA.
If you’re migrating from ElevenLabs, the flush semantics are the biggest behavioral difference. See the ElevenLabs migration guide.

Next steps

Latency

How to measure TTFA correctly on your deployment

Turn lifecycle

Flush semantics, the 5 s idle auto-flush, session reuse
Last modified on September 28, 2026