Skip to main content
KugelAudio normalizes request text so it is spoken naturally: numbers, dates, times, and currencies become words before synthesis.

Text Normalization

Text normalization converts numbers, dates, times, and other non-verbal text into spoken words:
  • “I have 3 apples” → “I have three apples”
  • “The meeting is at 2:30 PM” → “The meeting is at two thirty PM”
  • “€50.99” → “fifty euros and ninety-nine cents”
Normalization is on by default (normalize=True in Python, normalize: true in JavaScript and JSON). It runs in the request language. If you omit language, the voice’s primary language (the first entry of its supported_languages) is used, and English if the voice has none. The language is never detected from the text, so German text sent to an English-primary voice without language is normalized as English.
Always set language when the text is not in the voice’s primary language. Otherwise numbers, dates, and currencies are read with the wrong language’s words.
normalize: false does not turn normalization off: it skips the language-model pass, and basic expansion of numbers, dates, times, and currencies still runs on every request.

Supported Languages

Spell Tags

Content inside <spell> tags is read character by character and bypasses normalization; the surrounding text is still normalized. Syntax, grouping, and per-language symbol words are on Spell tags.

Custom Pronunciation Dictionaries

When normalization and <spell> tags aren’t enough (brand names, product names, acronyms the model gets wrong), use a per-project dictionary. It applies to requests that send the dictionary’s project_id, and a change takes effect on the next request. Manage dictionaries from the dashboard or the API:

Next Steps

Generate Speech

Basic speech generation

Spell Tags

Read codes and emails character by character

Voice Agent Prompting

System prompt patterns for LLM-driven voice agents

Custom Dictionaries

Per-project pronunciation fixes for brand names and acronyms
Last modified on September 22, 2026