> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kugelaudio.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech enhancement

> Remove background noise from speech, or keep only one speaker's voice, with Clarity

Speech enhancement cleans up recorded or live speech with Clarity (`clarity-1`).
It removes background noise and keeps the speech, or, when you give it a short
sample of one person, it keeps only that person's voice and removes everything
else.

| Task | What it does | Use it when |
| - | - | - |
| `noise_removal` | Removes background noise and keeps the speech. | One person is talking over traffic, fans, keyboard clicks or room noise. |
| `target_speaker_extraction` ("keep one voice") | Keeps one person's voice and removes every other voice and sound. You identify that person with a 2–8 second sample of them speaking. | Several people are audible (a call center floor, a meeting, a TV in the background) and you only want one of them. |

The SDKs pick the task for you: pass a speaker sample and they keep that voice,
leave it out and they remove noise. The output is always mono 16-bit PCM at
24 kHz, the same length as your input.

<Warning>
  Speech enhancement is available in preview. It is not yet intended for
  production workloads, and the interface may still change.
</Warning>

## Choose how to send audio

| You have | Use | Reference |
| - | - | - |
| A recording, up to 300 seconds | `client.enhance.generate`, or `POST /v1/audio/enhance` | [Python](/sdks/python/speech-enhancement), [JavaScript](/sdks/javascript/speech-enhancement), [REST](/api-reference/endpoints/speech-enhancement#enhance-a-recording) |
| Live audio, enhanced as it arrives | `client.enhance.stream`, or the `/v1/audio/enhance/stream` WebSocket | [Python](/sdks/python/speech-enhancement#enhance-live-audio), [JavaScript](/sdks/javascript/speech-enhancement#enhance-live-audio), [WebSocket](/api-reference/endpoints/speech-enhancement#enhance-live-audio) |
| A voice agent built on LiveKit or Pipecat | The input filter, which cleans the user's audio before VAD and STT hear it | [LiveKit](/integrations/livekit#5-remove-background-noise-optional), [Pipecat](/integrations/pipecat#removing-background-noise) |

## Quick start

Remove the noise from a recording, then keep only one voice from another. Every
request needs an [API key](/api-reference/authentication).

<CodeGroup>
  ```python Python theme={null}
  from kugelaudio import KugelAudio, load_audio

  client = KugelAudio(api_key="YOUR_API_KEY")

  # Remove background noise
  result = client.enhance.generate_sync(load_audio("call.wav"), model="clarity-1")
  result.save("call-clean.wav")

  # Keep one voice: pass a 2-8 second sample of that person
  result = client.enhance.generate_sync(
      load_audio("meeting.wav"),
      model="clarity-1",
      speaker=load_audio("speaker.wav"),
  )
  result.save("one-voice.wav")
  ```

  ```typescript JavaScript theme={null}
  import { KugelAudio, loadAudio } from 'kugelaudio';

  const client = new KugelAudio({ apiKey: 'YOUR_API_KEY' });

  // Remove background noise
  const clean = await client.enhance.generate(await loadAudio('call.wav'), {
    model: 'clarity-1',
  });
  await clean.save('call-clean.wav'); // Node.js; use clean.toBlob() in the browser

  // Keep one voice: pass a 2-8 second sample of that person
  const oneVoice = await client.enhance.generate(await loadAudio('meeting.wav'), {
    model: 'clarity-1',
    speaker: await loadAudio('speaker.wav'),
  });
  await oneVoice.save('one-voice.wav');
  ```
</CodeGroup>

The input can be any WAV with 16-, 24- or 32-bit integer or 32-bit float
samples, mono or stereo, at 8 to 48 kHz. Real-time streaming, the async and
blocking variants and every option are in the SDK references for
[Python](/sdks/python/speech-enhancement) and
[JavaScript](/sdks/javascript/speech-enhancement).

## Free month and paid plan

Every organization gets Clarity free until 30 October 2026, or for one calendar
month counted from when speech enhancement is enabled for it, whichever is later.
The **Enhancement** page of the
[dashboard](https://kugelaudio.com/dashboard) counts the days down. The free
month covers Clarity only, not text-to-speech or speech-to-text.

| | Free month | Paid |
| - | - | - |
| Cost | Not billed. | Billed from your organization's credits per second of input audio; see [pricing](https://kugelaudio.com/pricing). |
| Limits | Fair use: currently 10 requests per minute and 2 requests or streams at a time. | Your organization's normal API rate limits. |
| Ends | 30 October 2026, or one calendar month after enhancement was enabled if that is later. Requests are then refused with `402 FREE_PERIOD_ENDED` until you switch to paid. | Does not end. |

An organization owner or admin switches to paid on the **Enhancement** page of
the dashboard. The switch is one way: there is no going back to the free month.
It applies to new requests within about a minute; a stream that is already open
keeps the plan it started with until it closes.

How paid usage is billed:

* Per second of input audio processed, rounded up to the next whole second, with
  a one-second minimum per request.
* A stream that disconnects early is billed for the audio processed up to that
  point.
* A request that fails with an error is not billed.

## Limits and errors

| Status | `error_code` | What it means | What to do |
| -: | - | - | - |
| `429` | `RATE_LIMITED` | Your plan's limit is reached. `Rate limit exceeded (N requests per minute)` or `Concurrent generation limit reached (N)`. | For requests per minute, wait the `Retry-After` seconds. The concurrency limit has no `Retry-After`: retry once one of your open requests or streams has finished. On the free month, switching to paid moves you to your organization's own limits. |
| `402` | `INSUFFICIENT_CREDITS` | Paid plan only: your organization's credits are spent. | Top up your credits. |
| `402` | `FREE_PERIOD_ENDED` | Free plan only: your free Clarity month is over (or your organization has none). | Switch to paid on the Enhancement page of the dashboard. |
| `503` | `MODEL_UNAVAILABLE` | Not a limit of yours: enhancement is at capacity. | Retry after the `Retry-After` seconds. |

The SDKs raise the same error classes as for text-to-speech:

* `RateLimitError` for `429`, with the wait in `retry_after` (Python) or
  `retryAfter` (JavaScript) when the server sent one.
* `InsufficientCreditsError` for both `402`s.
* `AuthenticationError` for a missing or rejected API key, or an organization
  without access to speech enhancement.
* `ConnectionError` (`KugelAudioConnectionError` in Python) for `503` and
  network failures. In the browser, a live stream refused while connecting is
  always a `ConnectionError`, because the browser does not reveal the reason; see
  the [JavaScript SDK](/sdks/javascript/speech-enhancement).

Every status and `error_code` the enhancement endpoints return, including
validation errors, is in the
[endpoint reference](/api-reference/endpoints/speech-enhancement#errors); the
shared error format is in [Error codes](/api-reference/errors).

## A good speaker sample

Keeping one voice only works as well as the sample you give it.

* **2–8 seconds** of speech. Shorter or longer samples are rejected.
* **That person alone.** No one else talking, not even briefly.
* **Little background noise.** A quiet room is best. If you only have a noisy
  recording of the person, run it through noise removal first and use the
  result as the sample.

## Next steps

<CardGroup cols={2}>
  <Card title="Python SDK" icon="python" href="/sdks/python/speech-enhancement">
    Every enhancement method, argument and result field
  </Card>

  <Card title="JavaScript SDK" icon="js" href="/sdks/javascript/speech-enhancement">
    The same for Node.js and the browser
  </Card>

  <Card title="LiveKit" icon="signal-stream" href="/integrations/livekit#5-remove-background-noise-optional">
    Clean the caller's audio in a LiveKit agent
  </Card>

  <Card title="API reference" icon="code" href="/api-reference/endpoints/speech-enhancement">
    The REST and WebSocket wire protocol
  </Card>
</CardGroup>
