Skip to main content
Stream audio chunks as they’re generated for lower latency. One request per turn; the socket is reusable for further requests. For token-by-token text input, use Stream Input.

Connection

Connect with your API key:

Request Message

Send a JSON message to start generation. Fields share the meaning and defaults of the Generate Speech parameters:
boolean
default:"false"
Enable word-level timestamp alignment. When enabled, a word_timestamps message is sent after the audio chunks with per-word timing data.
number
default:"1.0"
Playback speed multiplier. Range: 0.8 (20% slower) to 1.2 (20% faster). Uses pitch-preserving WSOLA.
integer[]
Per-request dictionary selection. With project_id, omission applies all active project dictionaries filtered by language; without project_id, none are loaded. [] opts out. A non-empty list requires project_id and applies exactly those project dictionaries (including inactive ones), bypassing the language filter. Also accepted in the config of /ws/tts/stream and /ws/tts/multi, where both fields are sticky for the session.
Text Normalization: Set normalize: true to convert numbers, dates, and symbols to spoken words. Set language when the text is not in the voice’s primary language (English if it has none); the language is not detected from the text.

Update Settings Message

This socket is reusable across requests. Send an update_settings message to set sticky generation-parameter defaults that fill any field a later request omits; a per-request value still wins. The server replies with settings_updated.
Only those six generation parameters are updatable; every field is optional. Identity, project, dictionary, and audio-format fields (voice_id, model_id, sample_rate, output_format, project_id, dictionary_ids) are not; include one and the message is rejected with a VALIDATION_ERROR frame (the socket stays open).

Cancel Message (barge-in)

Abandons the request that is currently generating: no further audio frames are emitted for it and no final; the server acknowledges with interrupted instead. The socket stays open, so the next request can be sent immediately. A cancel with nothing in flight is acknowledged the same way. See Barge-in.

Response Messages

Audio Chunk

Field-by-field reference: Audio formats.

Word Timestamps (when word_timestamps: true)

score is a compatibility field and is currently always 1.0.

Settings Updated

Acknowledges an update_settings message; settings holds all sticky generation-parameter defaults now in effect (here, after the update_settings message shown above):

Interrupted

Acknowledges a cancel. It replaces final for that request; a cancelled request never finalizes:

Final Message

On this endpoint, final is the request-complete message and carries the request’s stats and usage. (The streaming endpoints emit a lighter end-of-audio final without usage, followed by session_closed. See Turn lifecycle.)
The usage object reports what this request consumed and what it was charged, so you can bill your own customers per request:

Example

Errors

WebSocket error frames use the same JSON error shape as HTTP responses:
A refused connection (bad key, connection limit) is an HTTP status on the upgrade, not a close code. See Error Codes for the full lookup table.
Last modified on September 22, 2026