<spell> tags causes each character to be read out
individually. Useful for email addresses, verification codes, acronyms, and
serial numbers.
Content inside
<spell> bypasses text normalization; normalize: true
still normalizes the surrounding prose. Always set language so letters,
digits and symbols use that language’s spoken words.Character translations by language
Letters are read by their names and digits as number words (“one”, “zwei”). Symbols are read as:
Whitespace inside a spell block is read as the language’s own word where the
voices say it more reliably (French “espace”, Spanish “espacio”, Italian
“spazio”, Polish “spacja”) and as the English word “space” in the other
languages.
Accented letters are read by their spoken name in French, Spanish and
Portuguese (“E accent aigu”, “eñe”, “C cedilha”). Where a language has no
name for an accented letter, it is read as its base letter, so
José is
spelled “J, O, S, E” in English and Dutch.
A few letters use their native name where the voices otherwise fall back to
the English one: Spanish “zeta” and “i griega”, Italian “i lunga”, Polish
“igrek” and “fau”, Portuguese “agá” and “capa”, Dutch “jee”.
Other characters without a spoken word, such as _, are passed
through unchanged and count as a letter for grouping. Write the word yourself
outside the tag when it matters (“underscore”, “Unterstrich”).
Grouping
Codes are grouped automatically: a 500 ms pause (· below) every four
characters, the way a human reads a code aloud.
.,
-, @, /, …), and each run is grouped on its own. Only runs that
contain a digit are grouped automatically. Runs made only of letters, such as
names and the parts of an email address, are read without pauses, because a
pause inside a spelled word can make the voice repeat the letter next to it:
group="N". An explicit size groups every
run, letters included:
group="0" to switch grouping off and have the content read as one
unbroken run (A, four, B, nine, X, Z).
The pauses need a model with break support. On most
legacy model IDs the groups are separated by a
plain space instead of a pause. Use
kugel-3.Examples
- Python
- JavaScript
- cURL
Pitfalls
- No nesting. A
<spell>tag inside another spell block is read as literal characters. - No break tags inside spell blocks. Use grouping for pacing instead.
Spell tags in streaming
When streaming text token by token, a bare<spell> tag may be split across
messages: the server buffers text until the closing </spell> arrives before
generating audio, and auto-closes an unfinished tag if the stream ends. This
buffering does not cover <spell group="N">: send a grouped tag, from
opening to closing tag, within one message. See
Streaming overview.
When spelling isn’t enough
If a brand name or domain term is pronounced wrong (rather than needing to be spelled out), use a pronunciation dictionary instead. It rewrites or IPA-annotates the word without changing your request text.Next steps
Pronunciation & IPA
Fix how specific words are spoken
Breaks
Explicit pauses outside spelled content
Streaming
Send spell tags from an LLM token stream
Text normalization
How numbers, dates and currencies around the tag are read