client.enhance sends speech to Clarity (clarity-1) and returns it cleaned:
background noise removed, or, when you pass speaker=, only that person’s
voice kept. What the two tasks do, the free month, the paid plan and the limits
are in the Speech enhancement guide; the wire
protocol is in the API reference.
generate and stream are async; generate_sync and stream_sync are their
blocking twins and take the same arguments.
Enhance a recording
asyncio:
Keep one voice
Pass a 2–8 second sample of the person to keep asspeaker=. The SDK then
switches to target_speaker_extraction; see
A good speaker sample.
generate / generate_sync
Returns an
EnhancedAudio with the same duration as the input.
The request is bounded by the client’s timeout (default 60 seconds).
Enhance live audio
stream opens a WebSocket, sends your audio while you iterate, and yields
enhanced mono 16-bit PCM at 24 kHz as it is produced. Iteration ends once the
last input chunk has been enhanced; the total output matches the input’s
duration.
speaker=load_audio("speaker.wav") to keep one voice. stream_sync is the
blocking form; it sends from a background thread:
Raw PCM input
Audio that is not a WAV file, such as a live capture, goes in as any iterable (or, forstream, async iterable) of mono 16-bit little-endian PCM bytes
chunks, with its sample_rate:
stream / stream_sync
Breaking out of the loop closes the connection and stops sending. With
stream,
wrap the iterator in contextlib.aclosing to make that cleanup immediate. The
client’s timeout bounds the wait for the stream to become ready.
Loading audio
load_audio
generate and for speaker=.
The file is sent as-is and decoded by the server, so every supported WAV format
works. The returned Audio has:
load_audio_stream
stream. Stereo and multi-channel audio is mixed
down to mono, and the sample rate is read from the file. chunk_seconds (greater
than 0, at most 1) sets the length of each chunk; chunks are sent as fast as the
connection allows, and small ones let enhanced audio start coming back sooner.
The returned AudioStream has sample_rate (Hz) and duration (seconds), and
iterating it yields the bytes chunks; iterating again starts from the
beginning. For 24-bit or float WAVs, use generate instead or convert to 16-bit.
EnhancedAudio
What generate returns: mono 16-bit PCM at 24 kHz.
Errors
Raised by the SDK itself, not the server:
Raised from the server’s answer, all subclasses of
KugelAudioError with
status_code, error_code, request_id and retry_after:
ValidationError (the server rejected the input), RateLimitError,
InsufficientCreditsError, AuthenticationError and
KugelAudioConnectionError. What each means and what to do is in
Limits and errors.
stream refused while connecting reports both 402s with error_code
INSUFFICIENT_CREDITS; on the free month it means the month is over.
Next steps
- Speech enhancement guide: tasks, free month, paid plan and limits
- Speech enhancement API: REST and WebSocket wire protocol
- LiveKit and Pipecat: enhancement as a voice-agent input filter