clarity-1).
It removes background noise and keeps the speech, or, when you give it a short
sample of one person, it keeps only that person’s voice and removes everything
else.
The SDKs pick the task for you: pass a speaker sample and they keep that voice,
leave it out and they remove noise. The output is always mono 16-bit PCM at
24 kHz, the same length as your input.
Choose how to send audio
Quick start
Remove the noise from a recording, then keep only one voice from another. Every request needs an API key.Free month and paid plan
Every organization gets Clarity free until 30 October 2026, or for one calendar month counted from when speech enhancement is enabled for it, whichever is later. The Enhancement page of the dashboard counts the days down. The free month covers Clarity only, not text-to-speech or speech-to-text.
An organization owner or admin switches to paid on the Enhancement page of
the dashboard. The switch is one way: there is no going back to the free month.
It applies to new requests within about a minute; a stream that is already open
keeps the plan it started with until it closes.
How paid usage is billed:
- Per second of input audio processed, rounded up to the next whole second, with a one-second minimum per request.
- A stream that disconnects early is billed for the audio processed up to that point.
- A request that fails with an error is not billed.
Limits and errors
The SDKs raise the same error classes as for text-to-speech:
RateLimitErrorfor429, with the wait inretry_after(Python) orretryAfter(JavaScript) when the server sent one.InsufficientCreditsErrorfor both402s.AuthenticationErrorfor a missing or rejected API key, or an organization without access to speech enhancement.ConnectionError(KugelAudioConnectionErrorin Python) for503and network failures. In the browser, a live stream refused while connecting is always aConnectionError, because the browser does not reveal the reason; see the JavaScript SDK.
error_code the enhancement endpoints return, including
validation errors, is in the
endpoint reference; the
shared error format is in Error codes.
A good speaker sample
Keeping one voice only works as well as the sample you give it.- 2–8 seconds of speech. Shorter or longer samples are rejected.
- That person alone. No one else talking, not even briefly.
- Little background noise. A quiet room is best. If you only have a noisy recording of the person, run it through noise removal first and use the result as the sample.
Next steps
Python SDK
Every enhancement method, argument and result field
JavaScript SDK
The same for Node.js and the browser
LiveKit
Clean the caller’s audio in a LiveKit agent
API reference
The REST and WebSocket wire protocol