> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kugelaudio.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech Enhancement

> Remove background noise or keep one voice with the Python SDK

`client.enhance` sends speech to Clarity (`clarity-1`) and returns it cleaned:
background noise removed, or, when you pass `speaker=`, only that person's
voice kept. What the two tasks do, the free month, the paid plan and the limits
are in the [Speech enhancement guide](/features/speech-enhancement); the wire
protocol is in the [API reference](/api-reference/endpoints/speech-enhancement).

`generate` and `stream` are async; `generate_sync` and `stream_sync` are their
blocking twins and take the same arguments.

## Enhance a recording

```python theme={null}
import asyncio

from kugelaudio import KugelAudio, load_audio

client = KugelAudio(api_key="YOUR_API_KEY")

async def main():
    result = await client.enhance.generate(load_audio("call.wav"), model="clarity-1")
    print(result.duration, "seconds")
    result.save("call-clean.wav")

asyncio.run(main())
```

Without `asyncio`:

```python theme={null}
result = client.enhance.generate_sync(load_audio("call.wav"), model="clarity-1")
result.save("call-clean.wav")
```

### Keep one voice

Pass a 2–8 second sample of the person to keep as `speaker=`. The SDK then
switches to `target_speaker_extraction`; see
[A good speaker sample](/features/speech-enhancement#a-good-speaker-sample).

```python theme={null}
result = client.enhance.generate_sync(
    load_audio("meeting.wav"),
    model="clarity-1",
    speaker=load_audio("speaker.wav"),
)
result.save("one-voice.wav")
```

### `generate` / `generate_sync`

```python theme={null}
async def generate(audio: Audio, *, model: str, speaker: Audio | None = None) -> EnhancedAudio
def generate_sync(audio: Audio, *, model: str, speaker: Audio | None = None) -> EnhancedAudio
```

| Argument | Type | Description |
| - | - | - |
| `audio` | `Audio` | The recording, from [`load_audio`](#load_audio). Any WAV the API accepts: 16-, 24- or 32-bit integer or 32-bit float PCM, mono or stereo, 8–48 kHz, at most 300 seconds. |
| `model` | `str` | Required. `"clarity-1"`. |
| `speaker` | `Audio \| None` | A 2–8 second sample of one voice, from `load_audio`. Given: only that voice is kept. Omitted: noise is removed and every voice is kept. |

Returns an [`EnhancedAudio`](#enhancedaudio) with the same duration as the input.
The request is bounded by the client's `timeout` (default 60 seconds).

## Enhance live audio

`stream` opens a WebSocket, sends your audio while you iterate, and yields
enhanced mono 16-bit PCM at 24 kHz as it is produced. Iteration ends once the
last input chunk has been enhanced; the total output matches the input's
duration.

```python theme={null}
import asyncio

from kugelaudio import KugelAudio, load_audio_stream

client = KugelAudio(api_key="YOUR_API_KEY")

async def main():
    enhanced = bytearray()
    audio = load_audio_stream("call.wav")
    async for chunk in client.enhance.stream(audio, model="clarity-1"):
        enhanced += chunk  # mono PCM16 at 24 kHz, as it arrives

asyncio.run(main())
```

Add `speaker=load_audio("speaker.wav")` to keep one voice. `stream_sync` is the
blocking form; it sends from a background thread:

```python theme={null}
for chunk in client.enhance.stream_sync(load_audio_stream("call.wav"), model="clarity-1"):
    print(len(chunk), "bytes")
```

### Raw PCM input

Audio that is not a WAV file, such as a live capture, goes in as any iterable
(or, for `stream`, async iterable) of mono 16-bit little-endian PCM `bytes`
chunks, with its `sample_rate`:

```python theme={null}
def pcm_chunks(path, chunk_bytes=3200):  # 0.1 s of mono PCM16 at 16 kHz
    with open(path, "rb") as f:
        while chunk := f.read(chunk_bytes):
            yield chunk

for chunk in client.enhance.stream_sync(
    pcm_chunks("call.pcm"), model="clarity-1", sample_rate=16000
):
    print(len(chunk), "bytes")
```

Chunks of at most one second work best; longer ones are split before sending.

### `stream` / `stream_sync`

```python theme={null}
def stream(
    audio: AudioStream | Iterable[bytes] | AsyncIterable[bytes],
    *,
    model: str,
    sample_rate: int | None = None,
    speaker: Audio | None = None,
) -> AsyncIterator[bytes]

def stream_sync(
    audio: AudioStream | Iterable[bytes],
    *,
    model: str,
    sample_rate: int | None = None,
    speaker: Audio | None = None,
) -> Iterator[bytes]
```

| Argument | Type | Description |
| - | - | - |
| `audio` | `AudioStream`, or an iterable of `bytes` | An [`AudioStream`](#load_audio_stream) from `load_audio_stream`, or raw mono PCM16 chunks. `stream_sync` does not take async iterables. |
| `model` | `str` | Required. `"clarity-1"`. |
| `sample_rate` | `int \| None` | Sample rate of raw chunks, 8000–48000. Required for raw chunks; taken from an `AudioStream`, and must match it if you pass both. |
| `speaker` | `Audio \| None` | A 2–8 second sample of one voice, from `load_audio`. |

Breaking out of the loop closes the connection and stops sending. With `stream`,
wrap the iterator in `contextlib.aclosing` to make that cleanup immediate. The
client's `timeout` bounds the wait for the stream to become ready.

## Loading audio

### `load_audio`

```python theme={null}
def load_audio(source: str | os.PathLike | bytes) -> Audio
```

Loads a WAV file, from a path or its bytes, for `generate` and for `speaker=`.
The file is sent as-is and decoded by the server, so every supported WAV format
works. The returned `Audio` has:

| Field | Type | Description |
| - | - | - |
| `data` | `bytes` | The WAV file bytes. |
| `filename` | `str` | File name sent with the upload (`"audio.wav"` for bytes). |
| `duration` | `float \| None` | Duration in seconds when the header can be read locally, else `None` (for example float WAVs). |

### `load_audio_stream`

```python theme={null}
def load_audio_stream(source: str | os.PathLike | bytes, *, chunk_seconds: float = 0.1) -> AudioStream
```

Loads a **16-bit** PCM WAV for `stream`. Stereo and multi-channel audio is mixed
down to mono, and the sample rate is read from the file. `chunk_seconds` (greater
than 0, at most 1) sets the length of each chunk; chunks are sent as fast as the
connection allows, and small ones let enhanced audio start coming back sooner.

The returned `AudioStream` has `sample_rate` (Hz) and `duration` (seconds), and
iterating it yields the `bytes` chunks; iterating again starts from the
beginning. For 24-bit or float WAVs, use `generate` instead or convert to 16-bit.

## `EnhancedAudio`

What `generate` returns: mono 16-bit PCM at 24 kHz.

| Member | Type | Description |
| - | - | - |
| `audio` | `bytes` | Raw mono PCM16 samples, little-endian. |
| `sample_rate` | `int` | `24000`. |
| `duration` | `float` | Duration in seconds. |
| `wav` | `bytes` | The audio as WAV file bytes. |
| `save(path)` | method | Writes the audio to `path` as a WAV file. |

## Errors

Raised by the SDK itself, not the server:

| Exception | When |
| - | - |
| `TypeError` | `audio` (for `generate`) or `speaker` is not an `Audio`, for example a path string. Wrap it in `load_audio(...)`. |
| `ValidationError` | `model` is empty; `stream` got a path or WAV bytes (use `load_audio_stream`); raw chunks without `sample_rate`, or a `sample_rate` that does not match the `AudioStream`; a chunk that is not `bytes`; `load_audio_stream` got a WAV that is not 16-bit, or a `chunk_seconds` outside (0, 1]; an unreadable or empty file. |

Raised from the server's answer, all subclasses of `KugelAudioError` with
`status_code`, `error_code`, `request_id` and `retry_after`:
`ValidationError` (the server rejected the input), `RateLimitError`,
`InsufficientCreditsError`, `AuthenticationError` and
`KugelAudioConnectionError`. What each means and what to do is in
[Limits and errors](/features/speech-enhancement#limits-and-errors).

```python theme={null}
from kugelaudio import InsufficientCreditsError, RateLimitError

try:
    result = client.enhance.generate_sync(load_audio("call.wav"), model="clarity-1")
except RateLimitError as e:
    print("retry in", e.retry_after, "seconds")  # None for the concurrency limit
except InsufficientCreditsError as e:
    print(e.error_code, e.message)  # INSUFFICIENT_CREDITS or FREE_PERIOD_ENDED
```

A `stream` refused while connecting reports both `402`s with `error_code`
`INSUFFICIENT_CREDITS`; on the free month it means the month is over.

## Next steps

* [Speech enhancement guide](/features/speech-enhancement): tasks, free month, paid plan and limits
* [Speech enhancement API](/api-reference/endpoints/speech-enhancement): REST and WebSocket wire protocol
* [LiveKit](/integrations/livekit#5-remove-background-noise-optional) and [Pipecat](/integrations/pipecat#removing-background-noise): enhancement as a voice-agent input filter
