Skip to main content
On-premise deployment is an Enterprise feature. Contact sales to discuss your requirements, licensing, and deployment setup before following this guide.
This guide takes you from a supplied KugelAudio release to a running API, then through stopping, upgrading, and recovering it. Speech synthesis runs on your infrastructure. Your deployment’s API still calls the hosted KugelAudio API (api.kugelaudio.com) over outbound HTTPS for licensing, API-key authentication, usage reporting, and voice lookups, so an outage of that path affects request authentication, not only licensing. The main path below is a single-machine k3s installation, with one TTS model, an optional normalizer, and standalone Valkey. A Kubernetes administrator must prepare the host and GPU runtime first. For an existing multi-node installation, use its supplied hardware profile and replica counts.

1. Prepare

Get your deployment package

Your KugelAudio contact supplies the approved chart package, checksum, deployment license, registry credentials, and reviewed hardware configuration. Contact sales if you have not arranged your Enterprise deployment yet. If you already have an agreement, ask your KugelAudio contact for these files and credentials. Keep these in a working directory on the administration machine, one directory per version (for example kugelaudio-<approved-version>/), so a new release’s SHA256SUMS never overwrites the one you need for rollback:
The two YAML files contain deployment settings; credentials are stored separately. Retain each accepted version’s directory for recovery.

Prepare Kubernetes and the GPU

You need Linux, a supported Kubernetes/k3s version, Helm 3, kubectl, a working NVIDIA driver and container runtime, and the reviewed GPU device-plugin configuration. The application chart does not install these prerequisites. For a new k3s host, follow the K3s installation instructions using the version agreed for your deployment. Configure the NVIDIA runtime and GPU add-on using the hardware instructions supplied with your release before continuing.
Ready to continue: the intended node is Ready, the NVIDIA runtime works, and Kubernetes advertises the GPU resources required by your hardware profile. The bundled single-GPU profile uses shared GPU slots; slots do not provide separate GPU memory. On a k3s host, kubectl and Helm need an authorized kubeconfig. An administrator can run them from a root shell with KUBECONFIG=/etc/rancher/k3s/k3s.yaml; do not make that file world-readable. Allow DNS and outbound HTTPS for the configured KugelAudio control plane, container registries, and image downloads listed in your handover. Fully air-gapped operation is not supported. Use your customer CA bundle if an HTTPS proxy is present.

Inspect the package

Run the remaining shell examples in Bash from your working directory. Replace the package placeholder with the actual filename:
The package also contains INSTALL.md, README.md, and, from chart 1.0.2, OPERATIONS.md. Read the matching installation prerequisites with:

2. Configure

Create the deployment Secret

Create the namespace and save the chart’s Secret template to a file:
Replace the placeholders in secrets.yaml with these values, and set metadata.namespace to ${KUGELAUDIO_NAMESPACE} if you chose a namespace other than kugelaudio: Keep the Secret type kubernetes.io/dockerconfigjson and the default name kugelaudio-onprem-secrets. The optional browser workspace needs the additional credentials in the template. Apply the file, confirm the Secret exists without displaying its values, then delete the plaintext file:
If you manage secrets with SOPS, Sealed Secrets, or External Secrets, create the same Secret through that workflow instead of kubectl apply. Project API keys are created later for applications; they do not go into this Secret.

Set the model, capacity, and endpoint

Review the supplied values.customer.yaml and hardware.yaml together: File order matters: this guide applies customer values first and hardware.yaml last. The hardware file wins on overlapping fields, including concurrency and service exposure. Review the effective settings instead of assuming a change in the first file overrides the profile. Keep the approved GPU placement and memory settings. Changing a replica count does not add physical GPU capacity, and increasing the runtime limit does not increase a project’s API-key limit or license entitlement. The API Service is named ingress and serves port 8000. The single-GPU preset defaults to ClusterIP; use an existing customer gateway or explicitly configure supported exposure. LoadBalancer requires a working cluster load-balancer implementation. Keep a stable API URL for applications. A customer-managed proxy or gateway can preserve this URL when you switch deployments during an upgrade. Terminate TLS there when required by your network policy.

3. Install and verify

Installing the chart starts the application automatically:
Initial image downloads, model loading, and warm-up can take several minutes. The timeout is a failure deadline, not an expected startup duration.
Success: the release is deployed, application containers are Ready, and the Helm test passes. This test checks health, readiness, and model catalogue; verify actual speech in the next step before accepting traffic. Set the API URL assigned to your deployment:
Expect HTTP 200. If external routing is not ready, test locally with kubectl port-forward -n "${KUGELAUDIO_NAMESPACE}" svc/ingress 8000:8000 and use http://127.0.0.1:8000 from another terminal. Port forwarding is a temporary diagnostic, not the production endpoint.

4. Set up API access

Create a dedicated project API key in the KugelAudio Dashboard, following Getting your API key. This key authorizes generation; the deployment license and local operator key cannot replace it. In Bash, load the key without recording its value in shell history:
List voices available to this key:
Copy a voice’s numeric id into the example below. See voice selection and pagination if you need a different voice.
Open output.wav and confirm that you hear the supplied text. Conversion requires FFmpeg; the Python SDK can save WAV directly. See generation formats for raw audio. Configure your application’s SDK with this API key and the local base URL: Python uses api_url; JavaScript and Java use apiUrl. Store the key in server-side secret configuration and keep caller concurrency within the agreed limits.

5. Start, stop, and restart

Install starts the services. There is no separate application daemon to launch. The commands below apply to the single-machine profile with fixed replicas, one kugel-3 TTS Deployment, and standalone Valkey in a namespace dedicated to this release. If autoscaling or GitOps controls replicas, use that controller’s maintenance procedure; otherwise it may immediately undo a manual stop.

Stop

  1. Pause new requests at the caller or gateway.
  2. Let active generations and streaming connections finish, or close them deliberately within your agreed maintenance deadline.
  3. Scale the application Deployments down:
Stopped: all application Deployments have zero replicas and their running pods have terminated. Completed Helm-test pods can remain. GPU memory held by TTS and normalizer is released; standalone Valkey’s transient state is lost. The release and its Secret remain installed. Scaling is a temporary override. A later Helm update or start restores the replicas from the accepted values.
Stopping the k3s system service does not reliably stop application containers. See Stopping K3s. Stop the application as above before host maintenance.

Start after a stop

Reapply the currently accepted package and values, restoring the recorded replicas. This creates a new Helm revision without selecting a newer release:
Repeat installation verification and the speech test, then resume callers. If startup fails, leave callers paused and inspect the failing pod before retrying.

Restart one component

For an unhealthy TTS component or a configuration change requiring a restart:
Use deployment/ingress for the API or deployment/normalizer for the enabled normalizer. Restarting TTS on a single GPU interrupts generation until the replacement is ready. Verify health and speech before resuming callers.

6. Upgrade

This procedure updates the existing deployment on the same machine. Schedule a maintenance window: speech generation is unavailable while workloads are replaced and the model loads. The API URL stays the same.

Prepare the release

Obtain the approved new package and checksum into a new version directory. Read its changelog and review any required values changes; keep the old version’s directory with its package and accepted values. From the new directory:
Record the last accepted Helm revision as your rollback target. Confirm the current deployment can generate speech, the required images are reachable, and disk/GPU capacity is sufficient.

Apply the update

  1. Pause new requests and drain existing streams.
  2. Upgrade using the reviewed values:

Verify and resume

Keep callers paused and check the updated deployment:
Repeat the voice selection and speech request through your deployment’s API URL and listen to the result. Resume callers only when application containers are Ready, the Helm test and health request pass, and speech works with the voices and languages your application uses. After accepting the update, use its package for future starts:
Record that package path, its values files, and the new Helm revision in your deployment handover. Keep the previous package and values for rollback. Do not override individual image/weight versions or use --reuse-values. Helm 3’s --atomic rolls back a failed upgrade, but passing pod readiness does not prove speech works. If the speech test fails, use rollback while callers remain paused.

7. Roll back

Keep callers paused. Use the revision recorded before the update:
Repeat the verification checks and speech test, then resume callers. Reset your operations record to the restored package and values so the next start does not inadvertently reapply the rejected release. Rollback restores Helm-managed resources. It does not restore an externally changed Secret, customer CA, GPU driver, gateway configuration, or lost transient state. Old images and valid licensing must still be available.

Troubleshooting

Start with the failing workload:
For routine maintenance, monitor pod restarts, GPU and disk health, the API health endpoint, and an authenticated speech probe. Keep the chart, values, hardware profile, and matching operations manual available; back up credentials through your secret manager. A single machine has no node redundancy. For support, include the chart version, Helm revision, timestamp, failed check, and relevant reviewed logs. Exclude credential values, environment dumps, and signed URLs. Contact KugelAudio.

Next steps

Python SDK configuration

Point the SDK at your deployment’s API URL

Regions

Hosted endpoints, if you also use the cloud API
Last modified on September 24, 2026