The official Python SDK for KugelAudio provides a simple, Pythonic interface for text-to-speech generation with both synchronous and asynchronous support.
The source is MIT-licensed and on GitHub at
Kugelaudio/python-sdk; issues and pull requests
are welcome there.
Installation
The SDK requires Python 3.10 or newer.
Or with uv (recommended):
Quick Start
Create an API key at app.kugelaudio.com/settings/api-keys
and keep it out of your code, for example in an environment variable (see
Authentication). The SDK does not read the
environment on its own, so pass the key explicitly:
Pre-connecting for Low Latency
For latency-sensitive applications, pre-establish the WebSocket connection at startup to keep the handshake out of your first TTS request. See Latency.
In the examples below, play_audio stands for your own playback function; the SDK does not ship one.
Async Applications (Recommended)
Sync Applications
For synchronous code that needs connection reuse, keep a streaming session
open across sends:
The synchronous generate() and stream() wrappers create and close a
connection on their own event-loop thread, so client.connect() does not
pre-connect those calls. Use the async API or streaming_session_sync() when
connection reuse matters. See Latency for measurement guidance.
Complete Example
Local turn detection (optional)
The SDK also ships an optional, local end-of-turn detector for voice agents. It
runs on your machine, needs Python 3.11 or newer, and pulls in PyTorch,
Transformers and ONNX Runtime:
Load the detector once per process and create one session per conversation.
Supported languages are de, en, es, it and nl. pcm_chunk stands for
16 kHz mono PCM16 bytes of the user’s audio:
History requires a bundle trained for multiple context messages; single-message
bundles reject a nonempty history. Supply only messages already known at
scoring time and assistant text that was actually spoken. update_history()
replaces the history, reset_turn() keeps it, and reset_conversation()
clears it.
Next steps
-
Configuration: client options, lifecycle, region selection
-
Generate Speech: one-shot generation, parameters, word timestamps
-
Streaming: streaming sessions, barge-in, multi-context sessions
-
Text Normalization: languages and spell tags
-
Voices: list, create, and manage voices
-
Dictionaries: per-project pronunciation lists
-
Speech to Text: transcribe audio files
-
Types & Errors: exceptions, data models, enums
-
Agent skill: teach your coding assistant to write correct KugelAudio code
Last modified on September 23, 2026