Skip to main content
client.enhance sends speech to Clarity (clarity-1) and returns it cleaned: background noise removed, or, when you pass speaker=, only that person’s voice kept. What the two tasks do, the free month, the paid plan and the limits are in the Speech enhancement guide; the wire protocol is in the API reference. generate and stream are async; generate_sync and stream_sync are their blocking twins and take the same arguments.

Enhance a recording

Without asyncio:

Keep one voice

Pass a 2–8 second sample of the person to keep as speaker=. The SDK then switches to target_speaker_extraction; see A good speaker sample.

generate / generate_sync

Returns an EnhancedAudio with the same duration as the input. The request is bounded by the client’s timeout (default 60 seconds).

Enhance live audio

stream opens a WebSocket, sends your audio while you iterate, and yields enhanced mono 16-bit PCM at 24 kHz as it is produced. Iteration ends once the last input chunk has been enhanced; the total output matches the input’s duration.
Add speaker=load_audio("speaker.wav") to keep one voice. stream_sync is the blocking form; it sends from a background thread:

Raw PCM input

Audio that is not a WAV file, such as a live capture, goes in as any iterable (or, for stream, async iterable) of mono 16-bit little-endian PCM bytes chunks, with its sample_rate:
Chunks of at most one second work best; longer ones are split before sending.

stream / stream_sync

Breaking out of the loop closes the connection and stops sending. With stream, wrap the iterator in contextlib.aclosing to make that cleanup immediate. The client’s timeout bounds the wait for the stream to become ready.

Loading audio

load_audio

Loads a WAV file, from a path or its bytes, for generate and for speaker=. The file is sent as-is and decoded by the server, so every supported WAV format works. The returned Audio has:

load_audio_stream

Loads a 16-bit PCM WAV for stream. Stereo and multi-channel audio is mixed down to mono, and the sample rate is read from the file. chunk_seconds (greater than 0, at most 1) sets the length of each chunk; chunks are sent as fast as the connection allows, and small ones let enhanced audio start coming back sooner. The returned AudioStream has sample_rate (Hz) and duration (seconds), and iterating it yields the bytes chunks; iterating again starts from the beginning. For 24-bit or float WAVs, use generate instead or convert to 16-bit.

EnhancedAudio

What generate returns: mono 16-bit PCM at 24 kHz.

Errors

Raised by the SDK itself, not the server: Raised from the server’s answer, all subclasses of KugelAudioError with status_code, error_code, request_id and retry_after: ValidationError (the server rejected the input), RateLimitError, InsufficientCreditsError, AuthenticationError and KugelAudioConnectionError. What each means and what to do is in Limits and errors.
A stream refused while connecting reports both 402s with error_code INSUFFICIENT_CREDITS; on the free month it means the month is over.

Next steps

Last modified on September 29, 2026