> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kugelaudio.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech Enhancement

> Remove background noise or keep one voice with the JavaScript SDK

`client.enhance` sends speech to Clarity (`clarity-1`) and returns it cleaned:
background noise removed, or, when you pass `speaker`, only that person's voice
kept. What the two tasks do, the free month, the paid plan and the limits are in
the [Speech enhancement guide](/features/speech-enhancement); the wire protocol
is in the [API reference](/api-reference/endpoints/speech-enhancement).

It works in Node.js and in the browser. Only reading and writing files by path
needs Node.js; in the browser pass a `Blob`, `File`, `ArrayBuffer` or
`Uint8Array`.

## Enhance a recording

```typescript theme={null}
import { KugelAudio, loadAudio } from 'kugelaudio';

const client = new KugelAudio({ apiKey: 'YOUR_API_KEY' });

const result = await client.enhance.generate(await loadAudio('call.wav'), {
  model: 'clarity-1',
});
console.log(result.duration, 'seconds');
await result.save('call-clean.wav'); // Node.js only
```

In the browser, pass the file straight from an `<input type="file">` and play
the result:

```typescript theme={null}
const file = input.files![0];
const result = await client.enhance.generate(file, { model: 'clarity-1' });
audioElement.src = URL.createObjectURL(result.toBlob());
```

### Keep one voice

Pass a 2–8 second sample of the person to keep as `speaker`. The SDK then
switches to `target_speaker_extraction`; see
[A good speaker sample](/features/speech-enhancement#a-good-speaker-sample).

```typescript theme={null}
const result = await client.enhance.generate(await loadAudio('meeting.wav'), {
  model: 'clarity-1',
  speaker: await loadAudio('speaker.wav'),
});
await result.save('one-voice.wav');
```

### `generate`

```typescript theme={null}
generate(audio: AudioInput, options: EnhanceGenerateOptions): Promise<EnhancedAudio>
```

| Argument | Type | Description |
| - | - | - |
| `audio` | `AudioInput` | The recording: `await loadAudio(...)`, or the WAV file itself as a `Blob`, `File`, `ArrayBuffer` or `Uint8Array`. Any WAV the API accepts: 16-, 24- or 32-bit integer or 32-bit float PCM, mono or stereo, 8–48 kHz, at most 300 seconds. File paths are not accepted here; load them with `loadAudio`. |
| `options.model` | `string` | Required. `'clarity-1'`. |
| `options.speaker` | `AudioInput` | Optional. A 2–8 second sample of one voice. Given: only that voice is kept. Omitted: noise is removed and every voice is kept. |

Returns an [`EnhancedAudio`](#enhancedaudio) with the same duration as the input.

## Enhance live audio

`stream` opens a WebSocket, sends your audio in the background while you
iterate, and yields enhanced mono 16-bit PCM at 24 kHz as `Uint8Array` chunks
as they are produced. Iteration ends once the last input chunk has been
enhanced; the total output matches the input's duration.

```typescript theme={null}
import { KugelAudio, loadAudio, loadAudioStream } from 'kugelaudio';

const client = new KugelAudio({ apiKey: 'YOUR_API_KEY' });

const input = await loadAudioStream('call.wav');
for await (const chunk of client.enhance.stream(input, { model: 'clarity-1' })) {
  console.log(chunk.length, 'bytes'); // mono PCM16 at 24 kHz, as it arrives
}
```

Add `speaker: await loadAudio('speaker.wav')` to keep one voice.

### Raw PCM input

Audio that is not a WAV file, such as a live capture, goes in as any iterable or
async iterable of mono 16-bit little-endian PCM chunks (`Uint8Array` or
`ArrayBuffer`), with its `sampleRate`:

```typescript theme={null}
async function* capture(): AsyncGenerator<Uint8Array> {
  // yield mono PCM16 chunks at 16 kHz as your source produces them
}

for await (const chunk of client.enhance.stream(capture(), {
  model: 'clarity-1',
  sampleRate: 16000,
})) {
  play(chunk);
}
```

Chunks of at most one second work best; longer ones are split before sending.

### `stream`

```typescript theme={null}
stream(
  audio: AudioStream | Iterable<PcmChunk> | AsyncIterable<PcmChunk>,
  options: EnhanceStreamOptions,
): AsyncGenerator<Uint8Array, void, undefined>
```

| Argument | Type | Description |
| - | - | - |
| `audio` | `AudioStream`, or an iterable of `PcmChunk` | An [`AudioStream`](#loadaudiostream) from `loadAudioStream`, or raw mono PCM16 chunks. |
| `options.model` | `string` | Required. `'clarity-1'`. |
| `options.sampleRate` | `number` | Sample rate of raw chunks, 8000–48000. Required for raw chunks; taken from an `AudioStream`, and must match it if you pass both. |
| `options.speaker` | `AudioInput` | Optional. A 2–8 second sample of one voice. |

Breaking out of the loop closes the connection and stops sending. The client's
`timeout` bounds the wait for the stream to become ready.

## Loading audio

### `loadAudio`

```typescript theme={null}
loadAudio(source: AudioSource): Promise<LoadedAudio>
```

Loads a WAV file for `generate` and for `speaker`. `source` is a path (Node.js
only) or the WAV file as a `Blob`, `File`, `ArrayBuffer` or `Uint8Array`. The
file is sent as-is and decoded by the server, so every supported WAV format
works. The returned `LoadedAudio` has:

| Field | Type | Description |
| - | - | - |
| `data` | `Uint8Array` | The WAV file bytes. |
| `filename` | `string` | File name sent with the upload (`'audio.wav'` when there is none). |
| `duration` | `number \| undefined` | Duration in seconds when the header can be read locally. |

### `loadAudioStream`

```typescript theme={null}
loadAudioStream(source: AudioSource, options?: LoadAudioStreamOptions): Promise<AudioStream>
```

Loads a **16-bit** integer PCM WAV for `stream`. Stereo and multi-channel audio
is mixed down to mono, and the sample rate is read from the file.
`options.chunkSeconds` (greater than 0, at most 1, default `0.1`) sets the length
of each chunk; chunks are sent as fast as the connection allows, and small ones
let enhanced audio start coming back sooner.

The returned `AudioStream` has `sampleRate` (Hz) and `duration` (seconds), and
iterating it (with `for` or `for await`) yields `Uint8Array` chunks; iterating
again starts from the beginning. For 24-bit or float WAVs, use `generate`
instead or convert to 16-bit.

## `EnhancedAudio`

What `generate` returns: mono 16-bit PCM at 24 kHz.

| Member | Type | Description |
| - | - | - |
| `audio` | `Uint8Array` | Raw mono PCM16 samples, little-endian. |
| `sampleRate` | `number` | `24000`. |
| `duration` | `number` | Duration in seconds. |
| `wav` | `Uint8Array` | The audio as WAV file bytes. |
| `toBlob()` | `Blob` | The audio as an `audio/wav` Blob, for example for an `<audio>` element. |
| `save(path)` | `Promise<void>` | Writes the audio to `path` as a WAV file. Node.js only. |

The package also exports `ENHANCED_SAMPLE_RATE` (`24000`),
`TASK_NOISE_REMOVAL`, `TASK_TARGET_SPEAKER_EXTRACTION` and the types
`AudioInput`, `PcmChunk`, `AudioSource`, `EnhanceGenerateOptions`,
`EnhanceStreamOptions` and `LoadAudioStreamOptions`.

## Errors

Thrown by the SDK itself, not the server:

| Error | When |
| - | - |
| `TypeError` | `audio` (for `generate`) or `speaker` is not audio, for example a path string. Wrap it in `await loadAudio(...)`. |
| `ValidationError` | `model` is empty; `stream` got a WAV file or `LoadedAudio` (use `loadAudioStream`); raw chunks without `sampleRate`, or a `sampleRate` that does not match the `AudioStream`; a chunk that is not binary; `loadAudioStream` got a WAV that is not 16-bit integer PCM, or a `chunkSeconds` outside (0, 1]; an unreadable or empty file. |
| `KugelAudioError` | A file path was used outside Node.js. |

Thrown from the server's answer, all subclasses of `KugelAudioError` with
`statusCode`, `errorCode`, `requestId` and `retryAfter`: `ValidationError` (the
server rejected the input), `RateLimitError`, `InsufficientCreditsError`,
`AuthenticationError` and `ConnectionError`. What each means and what to do is
in [Limits and errors](/features/speech-enhancement#limits-and-errors).

```typescript theme={null}
import { InsufficientCreditsError, RateLimitError } from 'kugelaudio';

try {
  await client.enhance.generate(await loadAudio('call.wav'), { model: 'clarity-1' });
} catch (e) {
  if (e instanceof RateLimitError) {
    console.log('retry in', e.retryAfter, 'seconds'); // undefined for the concurrency limit
  } else if (e instanceof InsufficientCreditsError) {
    console.log(e.errorCode, e.message); // INSUFFICIENT_CREDITS or FREE_PERIOD_ENDED
  } else {
    throw e;
  }
}
```

In Node.js, a `stream` refused while connecting reports both `402`s with
`errorCode` `INSUFFICIENT_CREDITS`; on the free month it means the month is
over. A browser's `WebSocket` does not reveal why an upgrade was refused, so in
the browser every refusal while connecting (a bad key, a limit, the free month,
capacity) arrives as `ConnectionError`. Call `generate` once to see the reason.

## Next steps

* [Speech enhancement guide](/features/speech-enhancement): tasks, free month, paid plan and limits
* [Speech enhancement API](/api-reference/endpoints/speech-enhancement): REST and WebSocket wire protocol
