Skip to main content
KugelAudio has a built-in Cognigy Voice Gateway custom-vendor endpoint — no proxy server needed. Register KugelAudio once as a speech service in Cognigy’s Self-Service Portal, then select it in your flows.
Cognigy’s built-in ElevenLabs provider is hosted-only and has no custom base-URL field, so it cannot be pointed at KugelAudio. Add KugelAudio as a custom speech vendor instead, as described below.
There are two ways to register KugelAudio, and they are interchangeable — pick one:

Setup

1. Get your KugelAudio API key and a voice ID

  • API key — open the KugelAudio dashboard, then go to Settings → API Keys.
  • Voice ID — open the KugelAudio dashboard, go to Voices, pick a voice, and copy its ID. The numeric ID or the voice handle both work.

2. Add KugelAudio as a speech service

In the Cognigy Voice Gateway Self-Service Portal, go to Speech → Add speech service and choose a custom vendor.
1

Name the vendor

Pick any name, for example kugelaudio. You’ll use this name to select the vendor in your flows.
2

Set the TTS HTTP URL

Cognigy appends /synthesize/<vendor-name> to this URL automatically — enter the base URL exactly as shown, without a trailing path.
3

Set the Authentication Token

Paste your KugelAudio API key. Cognigy sends it as an Authorization: Bearer header, which KugelAudio authenticates on every request.
4

Set the voice

Enter the voice ID from step 1.
5

Enable text-to-speech

Turn on Use for text-to-speech. Leave Use for speech-to-text off — this endpoint provides TTS only.
6

Enable streaming (recommended)

Turn on Enable text-to-speech streaming so audio starts playing while it is still being generated, instead of after the whole utterance is synthesized.
7

Choose the account scope

Select which Cognigy accounts may use this vendor, then save.

3. Select the vendor in your flow

Registering the vendor does not switch your flow over on its own. Set the vendor name in the Custom parameter of the relevant nodes — Set Session Config, Say, Question, Optional Question, or Session Speech Parameters Config.
If you skip this step the flow keeps using whichever provider it used before, and nothing appears to change.

Verify your API key first

If synthesis fails, check the key on its own before changing anything in Cognigy:
200 means the key is good. 401 Invalid API key means the key is wrong, truncated, or from a different environment — a KugelAudio project key looks like sk-kug-proj_lu_… and is long, so a partial copy-paste is the usual cause.

How it works

Cognigy sends one POST per utterance:
encoding and sample_rate are sent only when TTS streaming is enabled on the vendor, and they select how the audio comes back: Language tags are BCP-47 (de-DE, en-US). KugelAudio uses the language part and ignores the region, so de-DE and de-AT both select German. See Voices for the languages each voice supports.

Configure voice style and text normalization

Cognigy does not forward arbitrary option objects to custom TTS providers. To configure KugelAudio per request, put the optional settings in Cognigy’s Custom (Voice) field:
Existing configurations that contain only a voice ID or handle continue to work unchanged. In configured form, voice is required exactly once and all settings can appear in any order. The setting names and supported values are case-insensitive. style defaults to the existing natural behavior and normalize defaults to true; unknown, duplicate, or malformed options return a 400 VALIDATION_ERROR instead of silently falling back. See Text processing for what normalization does.

Audio format

Supported sample rates are 8000, 16000, 22050, 24000, and 44100. Cognigy Voice Gateway requests 8000 for streaming telephony audio.

Alternative: register KugelAudio as an Azure Speech container

KugelAudio’s ingress also answers the Azure Speech on-premises container contract, so you can register it under Cognigy’s built-in Microsoft Azure Speech Services vendor instead of as a custom vendor. Nothing Azure runs; the URL points at KugelAudio.
1

Choose the vendor

Speech → Add speech service → Vendor: Microsoft Azure Speech Services.
2

Switch to the container

Leave Use hosted Azure service off and turn on Use Azure Docker container (on-prem).
3

Set the container URLs

Enter the base URL exactly as shown, with no path after it.Cognigy makes both fields mandatory even when speech-to-text is switched off. KugelAudio serves TTS at that URL; the same URL also answers the Azure recognition contract, but only as a stub that always returns NoMatch — see Speech-to-text below.
4

Set the subscription key

Paste your KugelAudio API key into Subscription key (if required). Cognigy sends it as Ocp-Apim-Subscription-Key, which KugelAudio authenticates on every request.
With streaming on, the Microsoft Speech SDK cannot set request headers on a WebSocket upgrade, so it sends the key in the query string instead. KugelAudio redacts it from its own logs, but an HTTP proxy or load balancer in front of KugelAudio may record the full URL — check its access-log configuration before going live.
5

Enable text-to-speech and streaming

Turn on Use for text-to-speech and Enable text-to-speech streaming. Leave Use for speech-to-text off.
6

Choose the voice

Cognigy’s Azure voice list is a fixed list of Microsoft voice names and cannot show KugelAudio voices, so enter a KugelAudio voice yourself — see Choosing a voice.

Choosing a voice

The voice name arrives inside the SSML that Cognigy generates. KugelAudio reads it in this order:
Leaving Cognigy’s stock Microsoft voice selected works, but every call is answered by the fallback voice. Set the voice explicitly.

Streaming

Both of Cognigy’s code paths are served at the same URL:

Audio formats

KugelAudio accepts Azure’s X-Microsoft-OutputFormat tokens (and the equivalent WebSocket setting) at the sample rates it supports — 8000, 16000, 22050, 24000 and 44100 Hz:
  • raw-*-16bit-mono-pcm and riff-*-16bit-mono-pcm
  • raw-8khz-8bit-mono-mulaw / -alaw and their riff- forms
  • audio-16khz-*-mono-mp3 and audio-24khz-*-mono-mp3
An unsupported token returns 400 rather than being resampled silently. Cognigy Voice Gateway asks for raw-8khz-16bit-mono-pcm or audio-16khz-32kbitrate-mono-mp3.

Speech-to-text

The Azure recognition endpoints under /v1/azure exist only so Cognigy’s mandatory Container URL for STT field validates. They always answer NoMatch and never transcribe. KugelAudio’s real speech-to-text is the speech-to-text API, which does not speak the Azure protocol. Use the custom-vendor route or the API directly for recognition.

Azure container troubleshooting

With streaming on, a failed synthesis closes the WebSocket with the error code and its message instead of returning empty audio, so the synthesis fails visibly rather than playing silence on the call. Every request, including failed ones, appears on the Logs page of your dashboard.

Limits

Each organization has a rate limit and a cap on concurrent generations. A live phone deployment can hold several calls open at once, so the concurrency cap is usually the one that matters — when it is exceeded, synthesis is rejected with 429 RATE_LIMITED (WebSocket close 4029 with streaming on). Check your organization’s limits in the dashboard before going live, and contact us if you need them raised for production traffic. See Error codes for the exact responses.

Regions

To keep traffic and data inside the EU, use the regional host instead:
See Regions for the full list of endpoints.

Troubleshooting

Every error response carries a readable message, which Cognigy records in its logs:
See Error codes for the full list.
Last modified on September 14, 2026