Text Normalization
Text normalization converts numbers, dates, times, and other non-verbal text into spoken words:- “I have 3 apples” → “I have three apples”
- “The meeting is at 2:30 PM” → “The meeting is at two thirty PM”
- “€50.99” → “fifty euros and ninety-nine cents”
normalize=True in Python, normalize: true
in JavaScript and JSON). It runs in the request language. If you omit
language, the voice’s primary language (the first entry of its
supported_languages) is used, and English if the voice has none. The
language is never detected from the text, so German text sent to an
English-primary voice without language is normalized as English.
- Python
- JavaScript
- Java
- cURL
normalize: false does not turn normalization off: it skips the
language-model pass, and basic expansion of numbers, dates, times, and
currencies still runs on every request.
Supported Languages
Spell Tags
Content inside<spell> tags is read character by character and bypasses
normalization; the surrounding text is still normalized. Syntax, grouping,
and per-language symbol words are on Spell tags.
Custom Pronunciation Dictionaries
When normalization and<spell> tags aren’t enough (brand names, product
names, acronyms the model gets wrong), use a per-project dictionary. It
applies to requests that send the dictionary’s project_id, and a change
takes effect on the next request.
Manage dictionaries from the dashboard or the API:
- Guide: Dictionaries, how pronunciation dictionaries work
- API:
/v1/dictionaries, full CRUD plus idempotent bulk replace - SDKs:
client.dictionaries.*(Python, JavaScript, Java)
Next Steps
Generate Speech
Basic speech generation
Spell Tags
Read codes and emails character by character
Voice Agent Prompting
System prompt patterns for LLM-driven voice agents
Custom Dictionaries
Per-project pronunciation fixes for brand names and acronyms