play_audio stands for your own playback
function; the SDK does not ship one.
Streaming Audio
Receive audio chunks as they are generated for lower latency:Async Streaming
For async applications:stream_async() reuses the client’s pooled WebSocket by default; pass
reuse_connection=False to open a fresh one.
LLM Integration: Streaming Sessions
For real-time TTS when streaming text from an LLM:Async Streaming Session
Synchronous Streaming Session
Session Reuse
End a session without closing the WebSocket to avoid reconnection overhead when starting a new session (see Turn lifecycle).end_session() and close() discard audio that is still in flight, so call
drain() first to receive the rest of the turn:
StreamingSessionSync has no
end_session() or update_config().
Barge-in (interrupt the current turn)
When the end user speaks over the agent, callcancel_current() to stop
generating the current turn immediately and drop any buffered or queued text,
without closing the WebSocket. Unlike end_session(), no remaining text
is synthesized; the turn is abandoned. The socket stays open so the next
send() starts the next turn right away.
cancel_current() returns once the server acknowledges, or after a short quiet
timeout if the server goes silent. Stop local playback as soon as you call it:
a few in-flight frames may arrive before the acknowledgement. See
Barge-in for the
full protocol. The synchronous wrapper exposes cancel_current() too.
Updating settings mid-session
Change generation parameters on a live connection without reconnecting viaupdate_settings(). It sends an explicit, acknowledged update and returns
the parameters now in effect:
cfg_scale, temperature, speed,
max_new_tokens, language, normalize. Identity and audio-format settings
(voice_id, model_id, sample_rate, output_format, project_id,
dictionary_ids) are fixed for the connection; change those with update_config() after
end_session(). The change applies to the next turn; call it between turns.
The server rejects an out-of-range value or a non-updatable field with a typed
error. StreamingSessionSync and MultiContextSession expose update_settings()
too (on the multi session it is session-scoped and applies to contexts started
after the update).
Streaming session reference
A session is created withstreaming_session(...) (async) or
streaming_session_sync(...) (sync). Both accept the same configuration:
voice_id, model_id, cfg_scale, temperature, max_new_tokens,
sample_rate, flush_timeout_ms, normalize, language, word_timestamps,
speed, dictionary_ids, project_id, and an on_word_timestamps callback.
Although voice_id defaults to None in the SDK signature, it must be set
before the first synthesis; omission fails with MISSING_VOICE_ID.
Dictionaries apply only when project_id is set (see
Dictionaries).
The async StreamingSession exposes:
StreamingSessionSync has a subset of this API without await/async for:
send(), flush(), and drain() return list[AudioChunk], and
cancel_current(), update_settings(), close(), and the
last_word_timestamps / last_final / last_usage properties behave the
same. It has no connect(), end_session(), or update_config(), so session
reuse and voice switching are async-only.
Tuning streaming latency
By default the server accumulates LLM tokens and only begins generating at natural sentence boundaries. Tune how eagerly it starts with these session parameters:flush_timeout_ms is a session factory argument. Set the other three with
session.update_config(...) before the first send:
Multi-Context Sessions
A multi-context session manages up to 20 independent audio-generation contexts over a single WebSocket (see limits). Each context has its own text buffer, voice settings, and generation queue: useful for multi-speaker conversations, pre-buffering one stream while another plays, or interleaving audio for dynamic dialogue.multi_context_session(...):
Every context that synthesizes text must have an effective voice. Set
default_voice_id on the session or pass voice_id to every
create_context() call; otherwise the first text sent to that context fails
with MISSING_VOICE_ID.
MultiContextSession methods:
Word Timestamps in Streaming
Word timestamps work with one-shot streams and streaming sessions. During a one-shot stream, they are yielded aslist[WordTimestamp] objects between
audio chunks:
Word Timestamps in Streaming Sessions
Request word-level time alignments alongside audio. Timestamps are delivered per chunk after the corresponding audio data:Next steps
- Types & Errors:
AudioChunk,StreamConfig,SessionUsage,WordTimestamp - Text Normalization: languages and spell tags in streaming