/ws/tts/stream is one logical TTS request per turn, regardless of how many
send calls you make. The server’s text buffer accumulates tokens and hands
each sentence to the model as soon as its closing ., ? or ! arrives.
Text with no sentence end yet is synthesized after 2 s without new input, so
end each sentence with punctuation or flush at the end of the reply. Inside a single
turn, model state (KV cache, voice conditioning) is preserved across chunks
so prosody stays natural.
On /ws/tts/stream, a flush ends the turn. The server synthesizes what
is buffered, sends final and then session_closed with its own usage
block, and the next text starts a new turn on a fresh engine session. There
is no mid-turn flush: flushing per sentence turns one reply into N turns, each
with its own model prefill (the full model time-to-first-audio, see
Latency), its own final/session_closed and its own usage
record. See Turn lifecycle.
Flush granularity: flush once per reply
What decides time-to-first-audio per segment is how often you flush, not how you split the text yousend. Raw LLM tokens sent without flush are fine:
the server’s buffer reassembles them and only hands sentence-sized work to
the model.
If the full text is available before TTS starts, send it in one message with
flush.
We deliberately don’t publish exact ms figures here: they depend on region,
voice, and deployment. To reproduce the comparison for your own deployment,
time each chunking strategy against your endpoint as described in
Measuring TTFA correctly.
Server-side chunking
Native streaming sessions use the server’s sentence-aware text buffer. Send tokens as they arrive and let that buffer decide when a complete unit is ready. Two message fields tune it (see the Stream Input reference):flush_timeout_ms(default500): the longest a period waits when only the next token can tell what it means: after a single letter (z.could becomez.B.) or after a number (1299.could become1299.99,3.could become3. Oktober); after a number it waits up to twice as long. The next token ends the wait. Every other sentence is synthesized as soon as it ends. Text with no sentence end yet waits 2 s.max_buffer_length: the most characters the buffer holds before it is synthesized regardless of boundaries.
chunk_length_schedule or
auto_mode are unknown to it: the server ignores them and replies with a
warning frame. The SDK code
for this pattern is in the
Streaming overview.
Handle backpressure
If audio arrives faster than you can play it, bound your buffer instead of letting it grow:Common mistakes
- Per-sentence
flush=true. Every flush ends the turn, so the next text pays a fresh prefill and produces its ownsession_closedand usage record. Flush once per reply. - One session per sentence. A new WebSocket handshake plus a fresh model prefill, every sentence. Keep the same session open for the whole assistant reply and reuse it for the next one. See Turn lifecycle.
- Client-side sentence buffering before
send. Unnecessary: the server already buffers tokens and chunks at sentence boundaries. Pre-buffering on the client just adds latency. - Calling
send(text, flush=true)per word “for lower latency.” It is the opposite: each flush is a separate turn and model call. Word-granular flushing produces the worst possible TTFA.
Next steps
Latency
How to measure TTFA correctly on your deployment
Turn lifecycle
Flush semantics, the 5 s idle auto-flush, session reuse