AgentSession as the tts provider and
keeps one WebSocket open per agent, so every turn streams over a warm connection.
1. Install
- Python
- TypeScript
livekit extra installs livekit-agents. The Deepgram, OpenAI and Silero
plugins are only needed for the STT, LLM and VAD in the example below; swap them
for the providers you use.2. Add it to your agent
Pass the plugin astts and call prewarm() while the call is being set up, so the
first response does not pay for the connection. Save the Python version as
voice_agent.py or the TypeScript version as voice_agent.ts.
- Python
- TypeScript
KUGELAUDIO_API_KEY, LIVEKIT_URL, LIVEKIT_API_KEY and
LIVEKIT_API_SECRET from the environment (full list under
Environment variables). Pick a voice in
Voices; the plugin has no default voice.
3. Tune the sound
speed, temperature and cfg_scale shape the audio. Set them in the constructor or
change them mid-call; either way the next turn picks up the new values.
- Python
- TypeScript
- Out-of-range values raise instead of being clamped:
speedoutside0.8to1.2ortemperatureoutside0.0to1.0is rejected before anything is applied, so a bad update leaves the current settings untouched. speedandtemperatureare session-wide. All turns share one socket and the server applies the last value it received. A change therefore takes effect for turns started after it; a turn already playing keeps the old value.
4. Word timestamps (optional)
Turn on the word-timestamps option (off by default, see Options) and the server aligns each audio chunk and delivers per-word timing alongside it. LiveKit uses this for barge-in and transcript sync. Timestamps arrive after their audio chunk, so playback never waits for them. Details in Word timestamps.5. Remove background noise (optional)
enhancement() puts KugelAudio speech enhancement
(Clarity, clarity-1) on the agent’s audio input, the slot LiveKit uses for
noise cancellation. It removes background noise before VAD and STT hear the
user. Pass speaker= with a clean 2 to 8 second sample of one person and it
keeps only that voice. Python only; it needs livekit-agents 1.3.10 or later.
On an older release, importing enhancement raises an ImportError that says
to upgrade:
room_input_options= accept the same processor
as room_io.RoomInputOptions(noise_cancellation=enhancement()). The input must be
mono, which is LiveKit’s default; on a multi-channel input enhancement turns
itself off with an error log.
Before you turn it on:
- Latency. Enhancement runs on KugelAudio’s servers, not on the device, so the audio makes a round trip over the network and the agent hears the user somewhat later than with an on-device filter. Frames keep their size; they arrive delayed.
- Plan and limits. Enhancement uses your organization’s Clarity plan: the free month, then paid. Every second of the participant’s audio is enhanced, silence included, and each running agent session holds one stream open, so on the free month its concurrency limit caps how many sessions can be enhanced at once. See Free month and paid plan.
- Connection. The processor opens its connection when the agent starts, before the participant speaks, and keeps it for the participant’s later tracks and after a sample-rate change. After 60 seconds without audio the server closes it and the next audio reconnects.
- Failures. While the stream connects, or reconnects after a rate limit or a
network error, the agent hears the original audio and a warning is logged. A
rejected API key, spent credits, an ended free month or an invalid request turn
enhancement off for the rest of the session with one error log; pass
on_error=to be told.
Reference
Options
- Python
- TypeScript
Keyword arguments of
kugelaudio.livekit.TTS(...); all of them except api_key,
region, base_url, http_session and sample_rate can be changed later with
update_options(...).The plugin can also be registered under LiveKit’s plugin namespace:
Sample rates
24000 Hz is the native rate and the right choice unless the transport needs
another one. 22050, 16000, 8000 and 44100 are resampled server-side; use
8000 or 16000 for telephony.
Environment variables
Using the plugin without AgentSession
Direct synthesize() and stream() calls
Direct synthesize() and stream() calls
AgentSession drives these for you. Call them directly only when you consume the
audio frames yourself.- Python
- TypeScript
Troubleshooting
MISSING_VOICE_ID
MISSING_VOICE_ID
The plugin has no default voice. Pass a voice ID in the constructor or set it at
runtime before the first turn. Browse voices in Voices.
API key not found
API key not found
Set
KUGELAUDIO_API_KEY in the environment or pass the API key option to the
constructor.First response is slow
First response is slow
- Call
prewarm()during call setup, not right before the first turn - Create one
TTSinstance per worker and reuse it; do not construct one per phrase - Measure from the same region as the API before changing synthesis parameters, see Latency
Speed or temperature change did not apply
Speed or temperature change did not apply
Both are session-wide: a change binds for turns started after it, and a turn
already playing finishes with the old value. Values outside the accepted range
raise instead of being clamped, so check for an exception at the call site.
WebSocket connection fails
WebSocket connection fails
Check the base URL option and that the API is reachable from the worker over
wss://. After aclose() (Python) or close() (TypeScript), create a new
instance for the next session.No word timestamps arrive
No word timestamps arrive
Word timestamps are off by default, so check that the option is on. If it is,
alignment failed server-side for that chunk: the server logs it and skips the
timestamps, and the audio is unaffected.
Next steps
Voice prompting
Prompt patterns that make an LLM sound good when spoken
Streaming best practices
Chunking, flushing and latency for LLM-driven agents
Barge-in
What happens to the current turn when the caller interrupts
Pipecat integration
The same plugin idea for Pipecat pipelines