Request Body
This is the canonical parameter reference for TTS generation. The WebSocket endpoints accept the same fields (plus their own session controls, see Stream Input).string
required
The text to convert to speech. It must contain at least 2 characters after
trimming whitespace and is limited to 10,000 characters. Supports inline
<break>, <spell>, and
<prosody rate> tags.
Other SSML has no stable stripping or interpretation contract; remove it
before sending text (see Prompting).
Shorter text or text over 10,000 characters returns 400 VALIDATION_ERROR;
text over the model’s or your plan’s per-request limit returns
413 VALIDATION_ERROR. Streaming chunks on the WebSocket endpoints are exempt
from the 2-character minimum.string
default:"kugel-3"
The model to use. Use
kugel-3 for new integrations. The legacy IDs listed on Models remain accepted for backwards compatibility. The legacy request field model is also accepted as an alias when model_id is omitted; do not send both. Unknown IDs return 400 VALIDATION_ERROR. Accepted model IDs are billed and shown in Dashboard usage as requested, even when they route through the current production model.integer | string
required
The numeric voice ID or the voice handle. Required: there is no default voice. A request without a
voice_id is rejected with 400 MISSING_VOICE_ID; a voice_id that doesn’t
exist (or isn’t visible to your API key) returns 404 NOT_FOUND.number
default:"2.0"
Classifier-free guidance scale. Range: 1.2-2.5 (inclusive); values outside this range are clamped into it. Higher values = more expressive.
number
Sampling variance (0.0 to 1.0). 0 = most stable, 1 = most variance. See
temperature guidance. Values outside the range
return
400 VALIDATION_ERROR. When omitted, uses the engine default,
consistently with the streaming endpoints.integer
default:"2048"
Maximum tokens to generate. Range: 1-2048. Limits output length. Values
outside the range return
400 VALIDATION_ERROR.integer
default:"24000"
Output sample rate in Hz. Options: 8000, 16000, 22050, 24000, 44100.Audio is generated natively at 24kHz. Other rates use server-side resampling.
Any other value returns
400 VALIDATION_ERROR.string
Combined codec + rate token (e.g.
ulaw_8000) for non-PCM output such as
G.711 telephony codecs. Opt-in; when set it is authoritative and must not
contradict an explicitly non-default sample_rate. The wire-default
sample_rate: 24000 is ignored when resolving the token. See
Audio formats.boolean
default:"true"
Enable text normalization (converts numbers, dates, etc. to spoken words).Set
language when the text is not in the voice’s primary language (English if it has none); the language is not detected from the text.string
ISO 639-1 language code for text normalization (e.g., ‘de’, ‘en’, ‘fr’).Supported: de, en, fr, es, it, pt, nl, pl, sv, da, no, fi, cs, hu, ro, el, uk, bg, tr, vi, ar, hi, zh, ja, ko, sk, sl, hr, sr, ru, he, fa, ur, bn, ta, yue, th, id, msIf not provided, the server uses the voice’s primary language (English if it has none). It does not detect the language from the text.
Other values return
400 VALIDATION_ERROR.boolean
default:"false"
WebSocket endpoints only. Enable word-level timestamp alignment. See
Word timestamps. Not accepted by this REST
endpoint: requests are strictly validated, so sending it here returns
400 Bad Request. Use a WebSocket endpoint or an SDK instead.number
default:"1.0"
Playback speed multiplier. Range:
0.8 (20% slower) to 1.2 (20% faster).Uses pitch-preserving time-stretching (WSOLA) so the voice pitch stays natural at any speed.
Applies to the whole request; wrap text in <prosody rate="slow|medium|fast|0.8-1.2"> to override
the rate for a span (see Speed).
Values outside the range return 400 VALIDATION_ERROR; unlike cfg_scale,
speed is not clamped.integer
Project whose custom dictionaries should be loaded. Omit it to generate
without project dictionaries. A non-empty
dictionary_ids selection
requires project_id; the API verifies that the caller can access that
project when dictionary_ids is non-empty. A bare inaccessible project_id
currently falls back to generation without a dictionary; an explicit
selection fails with 403 UNAUTHORIZED.integer[]
Per-request dictionary selection.
- Omitted: when
project_idis set, all active dictionaries of that project apply, filtered by language; withoutproject_id, no project dictionary is loaded. []: no dictionary applies to this request.[7, 9]: exactly those dictionaries apply, including inactive ones, bypassing the language filter.
project_id. IDs must belong to that project;
unknown IDs return a 400 before generation starts. Maximum 50 IDs.Temperature guidance
temperature controls how much the sampler varies across regenerations of the
same text. Lower values are closer to greedy decoding (stable, repeatable
reads); higher values are more expressive but less consistent.
Omit
temperature to use the engine default. Explicit values use the public
0.0 to 1.0 scale and override that default; a client preset that sends a value is
an explicit override.
Spell Tags
Wrap text in<spell> to read it letter by letter (emails, codes, acronyms).
Spell content bypasses normalization. See Spell.
Response
Returns the requested raw encoding as a streaming binary response. The default is PCM16 (audio/pcm); G.711 output uses audio/basic. For encoding details
and the watermark, see
Audio formats.
Response Headers:
With the default format, the response body is raw PCM 16-bit signed
little-endian audio data streamed as binary chunks.
Example
Errors
See Error Codes for the full TTS error lookup table, including HTTP status codes, WebSocket close codes, and rate-limit behavior.Related endpoints
Stream Speech
Same request, audio chunks streamed over a WebSocket
Stream Input
Token-by-token text input for LLM agents