client.enhance sends speech to Clarity (clarity-1) and returns it cleaned:
background noise removed, or, when you pass speaker, only that person’s voice
kept. What the two tasks do, the free month, the paid plan and the limits are in
the Speech enhancement guide; the wire protocol
is in the API reference.
It works in Node.js and in the browser. Only reading and writing files by path
needs Node.js; in the browser pass a Blob, File, ArrayBuffer or
Uint8Array.
Enhance a recording
<input type="file"> and play
the result:
Keep one voice
Pass a 2–8 second sample of the person to keep asspeaker. The SDK then
switches to target_speaker_extraction; see
A good speaker sample.
generate
Returns an
EnhancedAudio with the same duration as the input.
Enhance live audio
stream opens a WebSocket, sends your audio in the background while you
iterate, and yields enhanced mono 16-bit PCM at 24 kHz as Uint8Array chunks
as they are produced. Iteration ends once the last input chunk has been
enhanced; the total output matches the input’s duration.
speaker: await loadAudio('speaker.wav') to keep one voice.
Raw PCM input
Audio that is not a WAV file, such as a live capture, goes in as any iterable or async iterable of mono 16-bit little-endian PCM chunks (Uint8Array or
ArrayBuffer), with its sampleRate:
stream
Breaking out of the loop closes the connection and stops sending. The client’s
timeout bounds the wait for the stream to become ready.
Loading audio
loadAudio
generate and for speaker. source is a path (Node.js
only) or the WAV file as a Blob, File, ArrayBuffer or Uint8Array. The
file is sent as-is and decoded by the server, so every supported WAV format
works. The returned LoadedAudio has:
loadAudioStream
stream. Stereo and multi-channel audio
is mixed down to mono, and the sample rate is read from the file.
options.chunkSeconds (greater than 0, at most 1, default 0.1) sets the length
of each chunk; chunks are sent as fast as the connection allows, and small ones
let enhanced audio start coming back sooner.
The returned AudioStream has sampleRate (Hz) and duration (seconds), and
iterating it (with for or for await) yields Uint8Array chunks; iterating
again starts from the beginning. For 24-bit or float WAVs, use generate
instead or convert to 16-bit.
EnhancedAudio
What generate returns: mono 16-bit PCM at 24 kHz.
The package also exports
ENHANCED_SAMPLE_RATE (24000),
TASK_NOISE_REMOVAL, TASK_TARGET_SPEAKER_EXTRACTION and the types
AudioInput, PcmChunk, AudioSource, EnhanceGenerateOptions,
EnhanceStreamOptions and LoadAudioStreamOptions.
Errors
Thrown by the SDK itself, not the server:
Thrown from the server’s answer, all subclasses of
KugelAudioError with
statusCode, errorCode, requestId and retryAfter: ValidationError (the
server rejected the input), RateLimitError, InsufficientCreditsError,
AuthenticationError and ConnectionError. What each means and what to do is
in Limits and errors.
stream refused while connecting reports both 402s with
errorCode INSUFFICIENT_CREDITS; on the free month it means the month is
over. A browser’s WebSocket does not reveal why an upgrade was refused, so in
the browser every refusal while connecting (a bad key, a limit, the free month,
capacity) arrives as ConnectionError. Call generate once to see the reason.
Next steps
- Speech enhancement guide: tasks, free month, paid plan and limits
- Speech enhancement API: REST and WebSocket wire protocol