Skip to main content
Generate complete audio from text. This is the simplest way to get started: provide text and receive audio back.

Basic Generation

Generation Parameters

The parameters you’ll touch most often (Python/REST snake_case; JavaScript uses camelCase):
  • text (required) and model_id: use kugel-3
  • voice_id: the voice to speak with (Using voices)
  • cfg_scale: expressiveness (see the guide below)
  • temperature: sampling randomness (0 to 1); omit it to use the model default
  • normalize + language: text normalization. Always set the language when you know it; otherwise the voice’s primary language is used
  • speed: playback rate (see Speed control below)
  • word_timestamps: word-level timestamps. SDK and WebSocket only; POST /v1/tts/generate rejects it with 400
The Generate Speech API reference is the canonical parameter table: every field with type, default, range, and error behavior, including output_format, project_id, and dictionary_ids.

CFG Scale Guide

The cfg_scale parameter controls how closely the model follows the voice characteristics. Accepted range: 1.2–2.5 (inclusive). Values outside this range are clamped into it.

Speed Control

speed changes the playback rate of the whole request with pitch-preserving time-stretching. The range is 0.8 to 1.2 (default 1.0); values outside it are rejected (400 on REST). To change the rate for part of a request, wrap that text in <prosody rate="...">. Examples and rules: Speed and per-span speed. For pauses, codes, and pronunciation fixes, see the Prompting guide: <break> tags, <spell> tags, and the unsupported-tags table.

Example with Common Options

Async Generation

Playing Audio in the Browser

The JavaScript SDK provides utility functions for audio playback:

Pre-connecting for Low Latency

For latency-sensitive applications, open the WebSocket connection at application startup so the handshake is not part of your first request. The per-SDK calls are in Latency: Pre-connect at startup.

Word Timestamps

Per-word time alignments (for subtitles, karaoke, lip-sync, and barge-in) are available from the SDKs and the WebSocket endpoints, not from POST /v1/tts/generate. See Word timestamps.

Next Steps

Streaming

Lower latency with real-time audio streaming

Text Processing

Text normalization and spell tags

Voices

Browse and use different voices

Models

Learn about available models
Last modified on September 22, 2026