output_format codecs, the audio chunk wire format,
and the AI-generated audio marking (watermark and disclosure header).
Default format
- Encoding: PCM 16-bit signed little-endian (
pcm_s16le) - Channels: Mono (1 channel)
- Sample rate: 24000 Hz (default; native generation rate)
- Byte order: Little-endian
sample_rate values (8000, 16000, 22050, and 44100) use
server-side resampling. The combined native output_format tokens below do
not include pcm_44100; 44100 is available through the legacy integer
sample_rate field (and through the ElevenLabs-compatible format dialect).
AI-generated audio marking (EU AI Act Art. 50): All generated audio is
watermarked in-band and every response carries a disclosure header. See
AI-generated audio marking below.
Output formats (output_format)
By default the API emits linear PCM16 at sample_rate. To request a different
codec — for example G.711 µ-law/a-law for telephony — send the combined
output_format token instead of (or in addition to) sample_rate. The token
carries codec and rate as one value, so impossible combinations like
“µ-law at 24 kHz” cannot be expressed.
Notes:
- Backwards compatible. Omitting
output_formatis identical to the default behavior — you getpcm_s16leframes. The strict checks below only apply to requests that sendoutput_format. - Conflicts are rejected. Sending both
output_formatand an explicitly non-defaultsample_ratethat disagrees with the token’s rate returns aVALIDATION_ERROR(HTTP 400 / WS error frame). The value24000is treated as the released SDKs’ serialized wire default rather than evidence of an explicit conflicting choice. Prefer sending onlyoutput_format, or matching values. - Sticky streaming config. On Stream Input
and Multi-Context, a format sent on an
ordinary config/context message persists until another valid ordinary
message changes it. The acknowledged
update_settingscommand accepts only generation parameters and rejectsoutput_formatandsample_rate. - G.711 frame semantics. For
ulaw_8000/alaw_8000, audio frames carryenc: "mulaw"/"alaw",sr: 8000, andsamplesequals the byte length (1 byte/sample). Decode with the standard G.711 tables (e.g. Pythonaudioop.ulaw2lin(payload, 2);audioopleft the standard library in Python 3.13, so install theaudioop-ltspackage there). On REST, the response usesContent-Type: audio/basicandX-Audio-Format: mulaw/alaw.
Telephony example (µ-law 8 kHz)
Audio chunk fields
Every WebSocket endpoint streams audio as JSON frames with these fields:AI-generated audio marking
Every audio stream KugelAudio produces is marked as AI-generated in two independent ways, as required by EU AI Act Article 50 (Regulation (EU) 2024/1689):- An in-band watermark embedded in the audio signal itself.
- A disclosure header on every audio response.
Disclosure header
Every HTTP audio response carries:
WebSocket endpoints return the same header
(
x-kugelaudio-ai-generated: true) in the connection handshake response. It
appears alongside the existing X-Sample-Rate and X-Audio-Format headers on
REST responses.
Watermark
Every output is watermarked in-band by the TTS engine before output encoding, so everyoutput_format (and proxy-side encodings such as MP3) carries it. The
mark is imperceptible and identifies the source platform or deployment, not an
individual request.
Detecting the watermark
customer_id) needs a few seconds of audio;
shorter clips report ai_generated with customer_id left unset.
Limitations
- Lossy encoding, added noise, and editing can weaken detection and attribution.
- The detector is trained on speech; a positive result on non-speech audio is unreliable.
Related
Generate Speech
The canonical request parameter reference
ElevenLabs-compatible output
MP3 and ElevenLabs-shaped responses via the proxy