client.tts.streamingSession() instead of client.tts.stream(). The session endpoint (/ws/tts/stream) keeps a persistent WebSocket connection and accumulates LLM tokens server-side, starting generation at natural sentence boundaries.
In the examples on this page, playAudio and stopLocalPlayback stand for
your own playback functions; the SDK does not ship them.
Why not flush per sentence?
Callingsend(token, true) on every sentence feels intuitive, but it actually increases latency:
- Each flush triggers a full model prefill (the fixed cost of loading context into the model).
- The server’s KV cache cannot be reused across separate flushes, so each segment is cold.
- Word-level flushing adds avoidable latency per sentence compared to letting the server batch. See Latency.
chunkLengthSchedule and autoMode.
Basic usage
Session Reuse
End a session without closing the WebSocket to avoid reconnection overhead (see Latency):Barge-in (interrupt the current turn)
When the end user speaks over the agent, callcancelCurrent() to stop
generating the current turn immediately and drop any buffered or queued text,
without closing the WebSocket. Unlike endSession(), no remaining text is
synthesized; the turn is abandoned. The socket stays open so you can send() the
next turn right away.
cancelCurrent() resolves once the server acknowledges (onInterrupted
fires), or after a short quiet timeout if the server goes silent. Stop local
playback as soon as you call it: a few in-flight frames may arrive before the
acknowledgement. See Barge-in
for the full protocol.
Updating settings mid-session
Change generation parameters on a live connection without reconnecting viaupdateSettings(). It sends an explicit, acknowledged update and resolves
with the parameters now in effect:
cfgScale, temperature, speed,
maxNewTokens, language, normalize. Identity and audio-format settings
(voiceId, modelId, sampleRate, outputFormat, projectId,
dictionaryIds) are fixed for the connection; change those with updateConfig() after endSession().
The change applies to the next turn; call it between turns. The promise
rejects if the server refuses an out-of-range value. MultiContextSession
exposes updateSettings() too (session-scoped; applies to contexts started
after the update).
Chunking presets
Session methods
streamingSession(config, callbacks) returns a StreamingSession.
voiceId is optional in the TypeScript declaration but must be set before the
first synthesis; omission fails with MISSING_VOICE_ID. Set projectId (and
optionally dictionaryIds) in StreamConfig to apply pronunciation
dictionaries (see Dictionaries).
There is no
drain() or lastFinal in the JavaScript SDK: endSession() and
close() already deliver the tail audio, and the end-of-turn stats arrive in
the onFinal callback.
Multi-Context Sessions
A multi-context session manages up to 20 independent audio-generation contexts over a single WebSocket. Each context has its own text buffer, voice, and generation queue: useful for multi-speaker conversations, pre-buffering one stream while another plays, or interleaving audio for dynamic dialogue.send(), closeContext() and close() return immediately. close() shuts
the socket at once and does not wait for audio, so close each context first and
wait for its onContextClosed:
createMultiContextSession(config?):
Unlike the Python SDK, the JavaScript multi-context config has no
modelId or
wordTimestamps option.
Every context that synthesizes text must have an effective voice. Configure
defaultVoiceId or pass voiceId to each createContext() call; a context
with neither fails on its first text with MISSING_VOICE_ID.
MultiContextSession methods and properties:
Audio arrives via the
onChunk callback as a MultiContextAudioChunk:
an AudioChunk plus a contextId field identifying its context.
Multi-context types
Next steps
- Types & Errors:
StreamConfig,StreamingSessionCallbacks,SessionUsage,AudioChunk - Text Normalization: languages and spell tags in streaming