Skip to main content
This page is the single reference for what the TTS endpoints emit: the default PCM encoding, the opt-in output_format codecs, the audio chunk wire format, and the AI-generated audio marking (watermark and disclosure header).

Default format

  • Encoding: PCM 16-bit signed little-endian (pcm_s16le)
  • Channels: Mono (1 channel)
  • Sample rate: 24000 Hz (default; native generation rate)
  • Byte order: Little-endian
Other supported sample_rate values (8000, 16000, 22050, and 44100) use server-side resampling. The combined native output_format tokens below do not include pcm_44100; 44100 is available through the legacy integer sample_rate field (and through the ElevenLabs-compatible format dialect).
AI-generated audio marking (EU AI Act Art. 50): All generated audio is watermarked in-band and every response carries a disclosure header. See AI-generated audio marking below.

Output formats (output_format)

By default the API emits linear PCM16 at sample_rate. To request a different codec — for example G.711 µ-law/a-law for telephony — send the combined output_format token instead of (or in addition to) sample_rate. The token carries codec and rate as one value, so impossible combinations like “µ-law at 24 kHz” cannot be expressed. Notes:
  • Backwards compatible. Omitting output_format is identical to the default behavior — you get pcm_s16le frames. The strict checks below only apply to requests that send output_format.
  • Conflicts are rejected. Sending both output_format and an explicitly non-default sample_rate that disagrees with the token’s rate returns a VALIDATION_ERROR (HTTP 400 / WS error frame). The value 24000 is treated as the released SDKs’ serialized wire default rather than evidence of an explicit conflicting choice. Prefer sending only output_format, or matching values.
  • Sticky streaming config. On Stream Input and Multi-Context, a format sent on an ordinary config/context message persists until another valid ordinary message changes it. The acknowledged update_settings command accepts only generation parameters and rejects output_format and sample_rate.
  • G.711 frame semantics. For ulaw_8000 / alaw_8000, audio frames carry enc: "mulaw" / "alaw", sr: 8000, and samples equals the byte length (1 byte/sample). Decode with the standard G.711 tables (e.g. Python audioop.ulaw2lin(payload, 2); audioop left the standard library in Python 3.13, so install the audioop-lts package there). On REST, the response uses Content-Type: audio/basic and X-Audio-Format: mulaw/alaw.

Telephony example (µ-law 8 kHz)

Audio chunk fields

Every WebSocket endpoint streams audio as JSON frames with these fields:

AI-generated audio marking

Every audio stream KugelAudio produces is marked as AI-generated in two independent ways, as required by EU AI Act Article 50 (Regulation (EU) 2024/1689):
  1. An in-band watermark embedded in the audio signal itself.
  2. A disclosure header on every audio response.
You do not need to enable anything: marking is mandatory and applied on every synthesis request, on every endpoint (native REST and WebSocket, ElevenLabs-compatible REST and WebSocket, and Vapi).

Disclosure header

Every HTTP audio response carries: WebSocket endpoints return the same header (x-kugelaudio-ai-generated: true) in the connection handshake response. It appears alongside the existing X-Sample-Rate and X-Audio-Format headers on REST responses.

Watermark

Every output is watermarked in-band by the TTS engine before output encoding, so every output_format (and proxy-side encodings such as MP3) carries it. The mark is imperceptible and identifies the source platform or deployment, not an individual request.

Detecting the watermark

The detector runs on numpy and works offline. Clips under one second are rejected. Recovering the source id (customer_id) needs a few seconds of audio; shorter clips report ai_generated with customer_id left unset.

Limitations

  • Lossy encoding, added noise, and editing can weaken detection and attribution.
  • The detector is trained on speech; a positive result on non-speech audio is unreliable.

Generate Speech

The canonical request parameter reference

ElevenLabs-compatible output

MP3 and ElevenLabs-shaped responses via the proxy
Last modified on September 22, 2026