Skip to main content
client.enhance sends speech to Clarity (clarity-1) and returns it cleaned: background noise removed, or, when you pass speaker, only that person’s voice kept. What the two tasks do, the free month, the paid plan and the limits are in the Speech enhancement guide; the wire protocol is in the API reference. It works in Node.js and in the browser. Only reading and writing files by path needs Node.js; in the browser pass a Blob, File, ArrayBuffer or Uint8Array.

Enhance a recording

In the browser, pass the file straight from an <input type="file"> and play the result:

Keep one voice

Pass a 2–8 second sample of the person to keep as speaker. The SDK then switches to target_speaker_extraction; see A good speaker sample.

generate

Returns an EnhancedAudio with the same duration as the input.

Enhance live audio

stream opens a WebSocket, sends your audio in the background while you iterate, and yields enhanced mono 16-bit PCM at 24 kHz as Uint8Array chunks as they are produced. Iteration ends once the last input chunk has been enhanced; the total output matches the input’s duration.
Add speaker: await loadAudio('speaker.wav') to keep one voice.

Raw PCM input

Audio that is not a WAV file, such as a live capture, goes in as any iterable or async iterable of mono 16-bit little-endian PCM chunks (Uint8Array or ArrayBuffer), with its sampleRate:
Chunks of at most one second work best; longer ones are split before sending.

stream

Breaking out of the loop closes the connection and stops sending. The client’s timeout bounds the wait for the stream to become ready.

Loading audio

loadAudio

Loads a WAV file for generate and for speaker. source is a path (Node.js only) or the WAV file as a Blob, File, ArrayBuffer or Uint8Array. The file is sent as-is and decoded by the server, so every supported WAV format works. The returned LoadedAudio has:

loadAudioStream

Loads a 16-bit integer PCM WAV for stream. Stereo and multi-channel audio is mixed down to mono, and the sample rate is read from the file. options.chunkSeconds (greater than 0, at most 1, default 0.1) sets the length of each chunk; chunks are sent as fast as the connection allows, and small ones let enhanced audio start coming back sooner. The returned AudioStream has sampleRate (Hz) and duration (seconds), and iterating it (with for or for await) yields Uint8Array chunks; iterating again starts from the beginning. For 24-bit or float WAVs, use generate instead or convert to 16-bit.

EnhancedAudio

What generate returns: mono 16-bit PCM at 24 kHz. The package also exports ENHANCED_SAMPLE_RATE (24000), TASK_NOISE_REMOVAL, TASK_TARGET_SPEAKER_EXTRACTION and the types AudioInput, PcmChunk, AudioSource, EnhanceGenerateOptions, EnhanceStreamOptions and LoadAudioStreamOptions.

Errors

Thrown by the SDK itself, not the server: Thrown from the server’s answer, all subclasses of KugelAudioError with statusCode, errorCode, requestId and retryAfter: ValidationError (the server rejected the input), RateLimitError, InsufficientCreditsError, AuthenticationError and ConnectionError. What each means and what to do is in Limits and errors.
In Node.js, a stream refused while connecting reports both 402s with errorCode INSUFFICIENT_CREDITS; on the free month it means the month is over. A browser’s WebSocket does not reveal why an upgrade was refused, so in the browser every refusal while connecting (a bad key, a limit, the free month, capacity) arrives as ConnectionError. Call generate once to see the reason.

Next steps

Last modified on September 29, 2026