How It Works
- Upload reference audio - Provide clean, representative speech
- Processing - Our AI analyzes the voice characteristics
- Voice created - Use your new voice in any TTS request
Requirements
Audio Quality
For best results, your reference audio should be:- Format: WAV, MP3, OGG, M4A, or FLAC
- File size: At most 50 MiB per reference
- Channels: Mono preferred
- Quality: Clean, no background noise
Content Guidelines
✅ Good audio:- Clear speech with natural pacing
- Single speaker only
- Minimal background noise
- Natural emotional range
- Free of filler words (um, uh, ah, hmm) unless you want them in the output
- Multiple speakers
- Background music
- Heavy reverb or echo
- Whispered or shouted speech
- Heavily compressed audio
- Recordings with frequent filler sounds or hesitations
- Long gaps or extended silence between sentences, unless you want the cloned voice to reproduce those pauses
Creating a Voice Clone
Via Dashboard
- Go to Dashboard → Voices → Create Voice
- Upload your reference audio
- Enter a name and description
- Click Create Voice
- Wait for processing to finish
Via SDK
- Python
- JavaScript
- cURL
Using Cloned Voices
Once created, use your cloned voice like any other:- Python
- JavaScript
- cURL
Best Practices
Optimizing Voice Quality
Use high-quality source audio
Use high-quality source audio
The quality of your cloned voice depends heavily on the source audio. Use professional recordings when possible.
Provide diverse samples
Provide diverse samples
Include a range of intonations, emotions, and sentence types in your reference audio for a more natural clone.
Adjust CFG scale
Adjust CFG scale
Experiment within the supported
cfg_scale range of 1.2 to 2.5.Remove filler sounds from samples
Remove filler sounds from samples
If your output contains unwanted “um”s, “ah”s, or hesitations, re-record or edit your reference audio to remove them. The model faithfully reproduces what it hears in the samples — clean input produces clean, controllable output. You can always add fillers via your text prompts later.
Troubleshooting
Managing Cloned Voices
List Your Voices
- Python
- JavaScript
- cURL
Update Voice
- Python
- JavaScript
- cURL
Delete Voice
- Python
- JavaScript
- cURL
Managing Reference Audio
You can add and remove reference audio files after creating a voice.List References
Add Reference
- Python
- JavaScript
Delete Reference
- Python
- JavaScript
Publishing Voices
Request that your voice be made public. It will be marked as pending verification until reviewed by an admin.- Python
- JavaScript
Generating Voice Samples
Trigger sample audio generation after uploading at least one reference:sample_s3_path and a signed sample_url.
AI Transparency & Watermarking
Voice-cloned output uses the same in-band watermark and HTTP disclosure header as every other synthesis response. See AI-generated audio marking for the wire details, detector example, and limitations.Privacy & Ethics
Guidelines
- Get consent - Always obtain permission before cloning someone’s voice
- Disclose synthetic speech - Be transparent when using cloned voices in public-facing contexts
- No impersonation - Don’t use cloned voices to deceive or defraud
- Respect rights - Don’t clone voices of public figures without authorization
Next Steps
Using Voices
Browse and use available voices
Generate Speech
Generate audio with your cloned voice
Models
Learn about available models