Skip to main content

Text-to-Speech

Basic Generation

Generate complete audio and receive it all at once:
The Java SDK takes numeric voice IDs (voiceId(int)). To use a voice handle, resolve it with GET /v1/voices/{handle} and pass its id.
Pronunciation dictionaries apply only when the request sets .projectId(long). See Dictionaries.

Streaming Audio

Receive audio chunks as they are generated for lower latency:

Text Normalization

Normalization (on by default) turns numbers, dates, and symbols into spoken words. Set .language(...) whenever you know it: without it the text is normalized in the voice’s primary language (English if it has none).
For supported languages see Text Processing. For spelling out emails, codes, and acronyms with <spell> tags see Spell.

Word Timestamps

Request word-level time alignments alongside audio for subtitle synchronization, lip-sync, or barge-in handling.

With Generate

With Streaming

Word timestamps add no extra audio latency. They arrive shortly after the corresponding audio chunk. See Latency for typical numbers.

Models

List Available Models

Error Handling

Next steps

Last modified on September 23, 2026