On-premise deployment is an Enterprise feature. Contact sales
to discuss your requirements, licensing, and deployment setup before following
this guide.
api.kugelaudio.com) over outbound HTTPS for licensing, API-key
authentication, usage reporting, and voice lookups, so an outage of that path
affects request authentication, not only licensing.
The main path below is a single-machine k3s installation, with one TTS model,
an optional normalizer, and standalone Valkey. A Kubernetes administrator must
prepare the host and GPU runtime first. For an existing multi-node installation,
use its supplied hardware profile and replica counts.
1. Prepare
Get your deployment package
Your KugelAudio contact supplies the approved chart package, checksum, deployment license, registry credentials, and reviewed hardware configuration. Contact sales if you have not arranged your Enterprise deployment yet. If you already have an agreement, ask your KugelAudio contact for these files and credentials. Keep these in a working directory on the administration machine, one directory per version (for examplekugelaudio-<approved-version>/), so a new release’s
SHA256SUMS never overwrites the one you need for rollback:
Prepare Kubernetes and the GPU
You need Linux, a supported Kubernetes/k3s version, Helm 3, kubectl, a working NVIDIA driver and container runtime, and the reviewed GPU device-plugin configuration. The application chart does not install these prerequisites. For a new k3s host, follow the K3s installation instructions using the version agreed for your deployment. Configure the NVIDIA runtime and GPU add-on using the hardware instructions supplied with your release before continuing.KUBECONFIG=/etc/rancher/k3s/k3s.yaml; do not make that file world-readable.
Allow DNS and outbound HTTPS for the configured KugelAudio control plane,
container registries, and image downloads listed in your handover. Fully
air-gapped operation is not supported. Use your customer CA bundle if an HTTPS
proxy is present.
Inspect the package
Run the remaining shell examples in Bash from your working directory. Replace the package placeholder with the actual filename:INSTALL.md, README.md, and, from chart 1.0.2,
OPERATIONS.md. Read the matching installation prerequisites with:
2. Configure
Create the deployment Secret
Create the namespace and save the chart’s Secret template to a file:secrets.yaml with these values, and set
metadata.namespace to ${KUGELAUDIO_NAMESPACE} if you chose a namespace
other than kugelaudio:
Keep the Secret type
kubernetes.io/dockerconfigjson and the default name
kugelaudio-onprem-secrets. The optional browser workspace needs the
additional credentials in the template. Apply the file, confirm the Secret
exists without displaying its values, then delete the plaintext file:
kubectl apply.
Project API keys are created later for applications; they do not go into this Secret.
Set the model, capacity, and endpoint
Review the suppliedvalues.customer.yaml and hardware.yaml together:
File order matters: this guide applies customer values first and
hardware.yaml last. The hardware file wins on overlapping fields, including
concurrency and service exposure. Review the effective settings instead of
assuming a change in the first file overrides the profile.
Keep the approved GPU placement and memory settings. Changing a replica count
does not add physical GPU capacity, and increasing the runtime limit does not
increase a project’s API-key limit or license entitlement.
The API Service is named ingress and serves port 8000. The single-GPU
preset defaults to ClusterIP; use an existing customer gateway or explicitly
configure supported exposure. LoadBalancer requires a working cluster
load-balancer implementation.
Keep a stable API URL for applications. A customer-managed proxy or gateway can
preserve this URL when you switch deployments during an upgrade. Terminate TLS
there when required by your network policy.
3. Install and verify
Installing the chart starts the application automatically:kubectl port-forward -n "${KUGELAUDIO_NAMESPACE}" svc/ingress 8000:8000
and use http://127.0.0.1:8000 from another terminal. Port forwarding is a
temporary diagnostic, not the production endpoint.
4. Set up API access
Create a dedicated project API key in the KugelAudio Dashboard, following Getting your API key. This key authorizes generation; the deployment license and local operator key cannot replace it. In Bash, load the key without recording its value in shell history:id into the example below. See
voice selection and pagination if
you need a different voice.
output.wav and confirm that you hear the supplied text. Conversion
requires FFmpeg; the Python SDK can save WAV directly.
See generation formats for raw audio.
Configure your application’s SDK with this API key and the local base URL:
Python uses api_url; JavaScript and
Java use apiUrl. Store the key in server-side secret
configuration and keep caller concurrency within the agreed limits.
5. Start, stop, and restart
Install starts the services. There is no separate application daemon to launch. The commands below apply to the single-machine profile with fixed replicas, onekugel-3 TTS Deployment, and standalone Valkey in a namespace dedicated to
this release. If autoscaling or GitOps controls replicas, use that controller’s
maintenance procedure; otherwise it may immediately undo a manual stop.
Stop
- Pause new requests at the caller or gateway.
- Let active generations and streaming connections finish, or close them deliberately within your agreed maintenance deadline.
- Scale the application Deployments down:
Stopping the k3s system service does not reliably stop application
containers. See Stopping K3s.
Stop the application as above before host maintenance.
Start after a stop
Reapply the currently accepted package and values, restoring the recorded replicas. This creates a new Helm revision without selecting a newer release:Restart one component
For an unhealthy TTS component or a configuration change requiring a restart:deployment/ingress for the API or deployment/normalizer for the
enabled normalizer. Restarting TTS on a single GPU interrupts generation until
the replacement is ready. Verify health and speech before resuming callers.
6. Upgrade
This procedure updates the existing deployment on the same machine. Schedule a maintenance window: speech generation is unavailable while workloads are replaced and the model loads. The API URL stays the same.Prepare the release
Obtain the approved new package and checksum into a new version directory. Read its changelog and review any required values changes; keep the old version’s directory with its package and accepted values. From the new directory:Apply the update
- Pause new requests and drain existing streams.
- Upgrade using the reviewed values:
Verify and resume
Keep callers paused and check the updated deployment:--reuse-values.
Helm 3’s --atomic rolls back a failed upgrade, but passing pod readiness does
not prove speech works. If the speech test fails, use rollback
while callers remain paused.
7. Roll back
Keep callers paused. Use the revision recorded before the update:Troubleshooting
Start with the failing workload:
For routine maintenance, monitor pod restarts, GPU and disk health, the API
health endpoint, and an authenticated speech probe. Keep the chart, values,
hardware profile, and matching operations manual available; back up credentials
through your secret manager. A single machine has no node redundancy.
For support, include the chart version, Helm revision, timestamp, failed check,
and relevant reviewed logs. Exclude credential values, environment dumps, and
signed URLs. Contact KugelAudio.
Next steps
Python SDK configuration
Point the SDK at your deployment’s API URL
Regions
Hosted endpoints, if you also use the cloud API