Skip to main content

Basic Generation

Generate complete audio and receive it all at once:
to_float32(), to_wav_bytes(), and save(..., format="wav") expect PCM16 output. For ulaw_8000 or alaw_8000, use audio.audio as raw G.711 bytes.

Generation parameters

What each parameter does, with defaults and ranges, is documented once in Generation parameters. The Python SDK takes them as snake_case keyword arguments on generate(), generate_async(), stream(), and stream_async(): text, model_id, voice_id, cfg_scale, temperature, max_new_tokens, sample_rate, output_format, normalize, language, word_timestamps, speed, dictionary_ids, project_id. Python-specific notes:
  • voice_id defaults to None in the signature, but synthesis fails with MISSING_VOICE_ID without it.
  • temperature=None leaves the value unset so the server applies its default.
  • Dictionaries apply only when project_id is set; dictionary_ids without it is rejected. See Dictionaries.

Async Generation

Word Timestamps with Generate

Request word-level time alignments alongside audio when using generate():
Word timestamps are also available with async generation:

Models

List Available Models

Next steps

Last modified on September 23, 2026