Skip to main content
KugelAudio provides an official TTS service for Pipecat, enabling high-quality voice synthesis in your voice AI pipelines.
Python only. Pipecat’s pipeline and TTSService framework is Python-only; there is no JavaScript/TypeScript equivalent to subclass, so the KugelAudio JS SDK does not ship a Pipecat service. Pipecat’s JS package (@pipecat-ai/client-js) is a browser client that connects to a Python Pipecat server over WebRTC/RTVI: run the KugelAudio Pipecat service (below) on that Python server and connect your JS/TS front-end to it. For a fully server-side JS/TS voice agent, use the LiveKit integration, which the JS SDK supports natively via kugelaudio/livekit.

Why Use KugelAudio with Pipecat?

  • Native service: Drop-in TTSService for Pipecat pipelines
  • Persistent WebSocket: Connection reuse keeps the handshake off the hot path
  • Built-in metrics: Automatic TTFB (Pipecat’s name for time-to-first-audio) and usage metrics
  • Low latency: streaming TTS built for real-time agents. See Latency for how to measure TTFA on your deployment.

Installation

This installs the KugelAudio SDK with its Pipecat dependency (pipecat-ai>=0.0.62 on Python 3.11 or newer). The integration supports both Pipecat 0.x and 1.x.

Quick Start

A complete local voice bot for Pipecat 1.x: microphone in, Deepgram STT, OpenAI LLM, KugelAudio TTS, speaker out. Install the Pipecat services it uses:
Save this as bot.py:
Set KUGELAUDIO_API_KEY, DEEPGRAM_API_KEY and OPENAI_API_KEY, then run python bot.py. Use headphones: the local audio transport has no echo cancellation, so speaker output can feed back into the microphone. For Pipecat 0.x, or another transport, STT or LLM, keep the tts lines and build the rest with the APIs of your Pipecat version.

Configuration

Service Parameters

Supported Sample Rates

Lower rates are resampled server-side. Any other value raises ValueError in the constructor.

Performance

Set language when the text is not in the voice’s primary language, and call prewarm() inside your async setup code (it needs a running event loop and does nothing otherwise) so the first reply does not pay the WebSocket handshake. The service then reuses that connection across run_tts() calls and reconnects transparently if it drops. On Pipecat 1.x each assistant turn still opens a new server-side context; the service creates it when the LLM starts responding, so the setup overlaps the LLM’s time-to-first-token. The KugelAudio TTFA: log line measures text send to first audio chunk on the WebSocket, without LLM or STT time. See Latency for how to measure end to end and Chunking & per-segment latency for how the server chunks text.

Usage Patterns

Updating the Voice at Runtime

Pipeline Frame Flow

The KugelAudioTTSService emits standard Pipecat frames, in this order:
  1. TTSStartedFrame: audio generation has begun
  2. TTSAudioRawFrame: raw PCM audio chunks (16-bit, mono), one or more
  3. TTSStoppedFrame: audio generation is complete
  4. ErrorFrame: if an error occurs during synthesis

Metrics Support

The service reports Pipecat’s TTFB metric (Pipecat’s name for time-to-first-audio, measured from request to first audio chunk) and TTS usage (characters per request) automatically; tts.can_generate_metrics() returns True.

Removing Background Noise

KugelAudioEnhanceFilter puts KugelAudio speech enhancement (Clarity, clarity-1) on the transport’s audio input, the audio_in_filter slot Pipecat uses for noise filters. It removes background noise before VAD and STT hear the user. Pass speaker= with a clean 2 to 8 second sample of one person and it keeps only that voice. The Quick Start bot with enhancement on the microphone input. Save it as bot.py, set the same three API keys, and run python bot.py:
Any Pipecat transport takes the filter the same way, as audio_in_filter= in its TransportParams. Push FilterEnableFrame(False) to switch it off mid-call (the stream closes and audio passes through) and FilterEnableFrame(True) to switch it back on. Before you turn it on:
  • Latency. Enhancement runs on KugelAudio’s servers, not on the device, so the audio makes a round trip over the network and the bot hears the user somewhat later than with an on-device filter. Frames keep their size; they arrive delayed.
  • Plan and limits. Enhancement uses your organization’s Clarity plan: the free month, then paid. Every second of input audio is enhanced, silence included, and each running pipeline holds one stream open, so on the free month its concurrency limit caps how many sessions can be enhanced at once. See Free month and paid plan.
  • Connection. The filter opens its connection when the pipeline starts, before the first audio, and closes it when the pipeline stops. After 60 seconds without audio the server closes it and the next audio reconnects.
  • Failures. While the stream connects, or reconnects after a rate limit or a network error, the bot hears the original audio and a warning is logged. A rejected API key, spent credits, an ended free month or an invalid request turn enhancement off for the rest of the session with one error log; pass on_error= to be told.

Environment Variables

Troubleshooting

Make sure KUGELAUDIO_API_KEY is set in your environment or pass api_key directly:
The service accepts 24000, 22050, 16000 and 8000. Set the same rate on the service and on your transport’s audio output (in the Quick Start, sample_rate=24000 and audio_out_sample_rate=24000).
Verify your base_url is correct and the KugelAudio API is reachable. The service connects via WebSocket (wss://) for audio streaming. If a persistent connection drops, the service automatically reconnects on the next run_tts() call.
Check Performance first: language unset and prewarm() not called (or called outside the event loop) are the usual causes. Then measure from the same region as the endpoint, see Measuring TTFA correctly.
The Pipecat integration requires Python 3.11 or newer. Check your version:

Next Steps

LiveKit Integration

Use KugelAudio with LiveKit Agents

Speech enhancement

Free month, paid plan and limits for the noise filter

Voice prompting

Prompt patterns that make an LLM sound good when spoken
Last modified on September 29, 2026