Skip to main content
Speech enhancement cleans up recorded or live speech with Clarity (clarity-1). It removes background noise and keeps the speech, or, when you give it a short sample of one person, it keeps only that person’s voice and removes everything else. The SDKs pick the task for you: pass a speaker sample and they keep that voice, leave it out and they remove noise. The output is always mono 16-bit PCM at 24 kHz, the same length as your input.
Speech enhancement is available in preview. It is not yet intended for production workloads, and the interface may still change.

Choose how to send audio

Quick start

Remove the noise from a recording, then keep only one voice from another. Every request needs an API key.
The input can be any WAV with 16-, 24- or 32-bit integer or 32-bit float samples, mono or stereo, at 8 to 48 kHz. Real-time streaming, the async and blocking variants and every option are in the SDK references for Python and JavaScript.

Free month and paid plan

Every organization gets Clarity free until 30 October 2026, or for one calendar month counted from when speech enhancement is enabled for it, whichever is later. The Enhancement page of the dashboard counts the days down. The free month covers Clarity only, not text-to-speech or speech-to-text. An organization owner or admin switches to paid on the Enhancement page of the dashboard. The switch is one way: there is no going back to the free month. It applies to new requests within about a minute; a stream that is already open keeps the plan it started with until it closes. How paid usage is billed:
  • Per second of input audio processed, rounded up to the next whole second, with a one-second minimum per request.
  • A stream that disconnects early is billed for the audio processed up to that point.
  • A request that fails with an error is not billed.

Limits and errors

The SDKs raise the same error classes as for text-to-speech:
  • RateLimitError for 429, with the wait in retry_after (Python) or retryAfter (JavaScript) when the server sent one.
  • InsufficientCreditsError for both 402s.
  • AuthenticationError for a missing or rejected API key, or an organization without access to speech enhancement.
  • ConnectionError (KugelAudioConnectionError in Python) for 503 and network failures. In the browser, a live stream refused while connecting is always a ConnectionError, because the browser does not reveal the reason; see the JavaScript SDK.
Every status and error_code the enhancement endpoints return, including validation errors, is in the endpoint reference; the shared error format is in Error codes.

A good speaker sample

Keeping one voice only works as well as the sample you give it.
  • 2–8 seconds of speech. Shorter or longer samples are rejected.
  • That person alone. No one else talking, not even briefly.
  • Little background noise. A quiet room is best. If you only have a noisy recording of the person, run it through noise removal first and use the result as the sample.

Next steps

Python SDK

Every enhancement method, argument and result field

JavaScript SDK

The same for Node.js and the browser

LiveKit

Clean the caller’s audio in a LiveKit agent

API reference

The REST and WebSocket wire protocol
Last modified on September 29, 2026