Python only. Pipecat’s pipeline and
TTSService framework is Python-only; there is no JavaScript/TypeScript equivalent to subclass, so the KugelAudio JS SDK does not ship a Pipecat service. Pipecat’s JS package (@pipecat-ai/client-js) is a browser client that connects to a Python Pipecat server over WebRTC/RTVI: run the KugelAudio Pipecat service (below) on that Python server and connect your JS/TS front-end to it. For a fully server-side JS/TS voice agent, use the LiveKit integration, which the JS SDK supports natively via kugelaudio/livekit.Why Use KugelAudio with Pipecat?
- Native service: Drop-in
TTSServicefor Pipecat pipelines - Persistent WebSocket: Connection reuse keeps the handshake off the hot path
- Built-in metrics: Automatic TTFB (Pipecat’s name for time-to-first-audio) and usage metrics
- Low latency: streaming TTS built for real-time agents. See Latency for how to measure TTFA on your deployment.
Installation
pipecat-ai>=0.0.62 on Python 3.11 or newer). The integration supports both
Pipecat 0.x and 1.x.
Quick Start
A complete local voice bot for Pipecat 1.x: microphone in, Deepgram STT, OpenAI LLM, KugelAudio TTS, speaker out. Install the Pipecat services it uses:bot.py:
KUGELAUDIO_API_KEY, DEEPGRAM_API_KEY and OPENAI_API_KEY, then run
python bot.py. Use headphones: the local audio transport has no echo
cancellation, so speaker output can feed back into the microphone. For
Pipecat 0.x, or another transport, STT or LLM, keep the tts lines and
build the rest with the APIs of your Pipecat version.
Configuration
Service Parameters
Supported Sample Rates
Lower rates are resampled server-side. Any other value raises
ValueError in
the constructor.
Performance
Setlanguage when the text is not in the voice’s primary language, and call
prewarm() inside your async setup code (it needs a running event loop and
does nothing otherwise) so the first reply does not pay the WebSocket
handshake. The service then reuses that connection across run_tts() calls
and reconnects transparently if it drops. On Pipecat 1.x each assistant turn
still opens a new server-side context; the service creates it when the LLM
starts responding, so the setup overlaps the LLM’s time-to-first-token. The
KugelAudio TTFA: log line measures text send to first audio chunk on the
WebSocket, without LLM or STT time. See Latency for how to measure
end to end and Chunking & per-segment latency
for how the server chunks text.
Usage Patterns
Updating the Voice at Runtime
Pipeline Frame Flow
TheKugelAudioTTSService emits standard Pipecat frames, in this order:
TTSStartedFrame: audio generation has begunTTSAudioRawFrame: raw PCM audio chunks (16-bit, mono), one or moreTTSStoppedFrame: audio generation is completeErrorFrame: if an error occurs during synthesis
Metrics Support
The service reports Pipecat’s TTFB metric (Pipecat’s name for time-to-first-audio, measured from request to first audio chunk) and TTS usage (characters per request) automatically;tts.can_generate_metrics() returns True.
Removing Background Noise
KugelAudioEnhanceFilter puts KugelAudio
speech enhancement (Clarity, clarity-1) on
the transport’s audio input, the audio_in_filter slot Pipecat uses for noise
filters. It removes background noise before VAD and STT hear the user. Pass
speaker= with a clean 2 to 8 second sample of one person and it keeps only
that voice.
The Quick Start bot with enhancement on the microphone input. Save it as
bot.py, set the same three API keys, and run python bot.py:
audio_in_filter= in its TransportParams. Push FilterEnableFrame(False)
to switch it off mid-call (the stream closes and audio passes through) and
FilterEnableFrame(True) to switch it back on.
Before you turn it on:
- Latency. Enhancement runs on KugelAudio’s servers, not on the device, so the audio makes a round trip over the network and the bot hears the user somewhat later than with an on-device filter. Frames keep their size; they arrive delayed.
- Plan and limits. Enhancement uses your organization’s Clarity plan: the free month, then paid. Every second of input audio is enhanced, silence included, and each running pipeline holds one stream open, so on the free month its concurrency limit caps how many sessions can be enhanced at once. See Free month and paid plan.
- Connection. The filter opens its connection when the pipeline starts, before the first audio, and closes it when the pipeline stops. After 60 seconds without audio the server closes it and the next audio reconnects.
- Failures. While the stream connects, or reconnects after a rate limit or a
network error, the bot hears the original audio and a warning is logged. A
rejected API key, spent credits, an ended free month or an invalid request turn
enhancement off for the rest of the session with one error log; pass
on_error=to be told.
Environment Variables
Troubleshooting
API key not found
API key not found
Make sure
KUGELAUDIO_API_KEY is set in your environment or pass api_key directly:Unsupported sample rate error
Unsupported sample rate error
The service accepts
24000, 22050, 16000 and 8000. Set the same rate on
the service and on your transport’s audio output (in the Quick Start,
sample_rate=24000 and audio_out_sample_rate=24000).WebSocket connection fails
WebSocket connection fails
Verify your
base_url is correct and the KugelAudio API is reachable. The service connects via WebSocket (wss://) for audio streaming. If a persistent connection drops, the service automatically reconnects on the next run_tts() call.High latency
High latency
Check Performance first:
language unset and prewarm() not
called (or called outside the event loop) are the usual causes. Then measure
from the same region as the endpoint, see
Measuring TTFA correctly.Python version incompatibility
Python version incompatibility
The Pipecat integration requires Python 3.11 or newer. Check your version:
Next Steps
LiveKit Integration
Use KugelAudio with LiveKit Agents
Speech enhancement
Free month, paid plan and limits for the noise filter
Voice prompting
Prompt patterns that make an LLM sound good when spoken