In the examples on this page, playAudio and playAudioChunk stand for your
own playback functions; the SDK does not ship them.
Basic Generation
Generate complete audio and receive it all at once:
Generation parameters
What each parameter does, with defaults and ranges, is documented once in
Generation parameters. The
JavaScript SDK takes them as camelCase fields of GenerateOptions:
text, modelId, voiceId, cfgScale, temperature, maxNewTokens,
sampleRate, outputFormat, normalize, language, wordTimestamps,
speed, projectId, dictionaryIds.
JavaScript-specific notes:
voiceId is optional in the type, but synthesis fails with
MISSING_VOICE_ID without it.
- When you set
outputFormat, also set the matching sampleRate (for
example outputFormat: 'pcm_16000' with sampleRate: 16000) so
audio.sampleRate describes the audio you received.
projectId and dictionaryIds select pronunciation dictionaries; see
Dictionaries.
AudioResponse carries no usage. For per-request usage, call stream() and
read stats.usage in onFinal.
Playing Audio in Browser
The WAV and PCM decoding utilities expect PCM16 input. For ulaw_8000 or
alaw_8000, consume audio.audio or decoded chunk bytes as raw G.711 data.
Streaming Audio
Receive audio chunks as they are generated for lower latency:
stream() reuses the client’s pooled WebSocket. Pass false as a third
argument (client.tts.stream(options, callbacks, false)) to open a fresh
connection for that call.
Streaming to a Node.js Readable
For server-side integrations that expect a Node.js Readable stream, such as Vapi custom TTS endpoints or Express/Fastify handlers, use client.tts.toReadable() instead of wiring onChunk manually.
toReadable() avoids a common race-condition: the stream object is returned before any audio arrives, so you can safely pipe() or attach listeners immediately.
toReadable() is Node.js only. It requires the built-in stream module and will throw in browser environments. Use the callback-based stream() API for browser code.
Call await KugelAudio.create(...) (or await client.connect()) at application startup. This pre-establishes the WebSocket connection so that subsequent toReadable() calls skip the connection overhead and start streaming audio immediately (see Latency). Like stream(), toReadable(options, false) opts out of the pooled connection.
Processing Audio Chunks
Word Timestamps
Request word-level time alignments alongside audio. Useful for subtitle synchronization, lip-sync, and barge-in handling.
With Generate
With Streaming
Word timestamps add no extra audio latency. They arrive shortly after the corresponding audio chunk. See Latency for typical numbers.
Models
List Available Models
Utility Functions
base64ToArrayBuffer
Convert base64 string to ArrayBuffer:
decodePCM16
Convert base64 PCM16 to Float32Array for Web Audio API:
createWavFile
Create a WAV file from PCM16 data:
createWavBlob
Create a playable Blob from PCM16 data:
client.tts.toReadable (Node.js only)
Convert a TTS stream directly to a Node.js Readable for use in HTTP handlers, pipelines, and server-side integrations. See Streaming to a Node.js Readable for a full example.
Next steps
Last modified on September 22, 2026