Basic Generation
Generate complete audio and receive it all at once:
to_float32(), to_wav_bytes(), and save(..., format="wav") expect PCM16
output. For ulaw_8000 or alaw_8000, use audio.audio as raw G.711 bytes.
Generation parameters
What each parameter does, with defaults and ranges, is documented once in
Generation parameters. The Python
SDK takes them as snake_case keyword arguments on generate(),
generate_async(), stream(), and stream_async():
text, model_id, voice_id, cfg_scale, temperature, max_new_tokens,
sample_rate, output_format, normalize, language, word_timestamps,
speed, dictionary_ids, project_id.
Python-specific notes:
voice_id defaults to None in the signature, but synthesis fails with
MISSING_VOICE_ID without it.
temperature=None leaves the value unset so the server applies its default.
- Dictionaries apply only when
project_id is set; dictionary_ids without
it is rejected. See
Dictionaries.
Async Generation
Word Timestamps with Generate
Request word-level time alignments alongside audio when using generate():
Word timestamps are also available with async generation:
Models
List Available Models
Next steps
Last modified on September 23, 2026