Use your own voice in your apps
Record yourself, name the voice, and speak with it through the same API call you use for everything else. English supported today.
Make it in three steps
1. Give us a voice
Record a paragraph in the browser, upload 10 to 60 seconds of one clear voice, or describe a voice with meters for gender, age, pitch, pace, warmth and energy.
2. Name it
Confirm you have the right to use the voice. It is private to your account and never available to other accounts.
3. Speak
Send its cv_ id as the voice in /v1/speak. That is the whole integration.
What happens behind the scenes
- Verification before use. A voice is marked ready only after we speak a test sentence with it, check it is real audio, and check with speech recognition that the words were heard. A voice that fails is removed.
- Starter lines. About two dozen common lines (greetings, thanks, holds, closings) are prepared in the background so they play fast.
- Designed voices are first rendered as a reference sample, then saved like a recording, so they sound the same every time.
- Limits: 3 voices per account. Replacing the sample bumps the voice version. Deleting a voice removes it at our provider and deletes every private recording made with it.
Languages (measured, not promised)
| Language | Status |
|---|---|
| English | Supported |
| French, German | Beta: understandable, with small errors |
| Spanish, Hindi, Chinese, Japanese | Not offered: not intelligible in our testing. Use a catalog voice |
Use it in your code
cURL
curl -X POST https://speakvora.com/api/v1/speak -H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
-d '{"text":"Thanks for calling.","voice":"cv_YOUR_VOICE_ID","lang":"en"}'
# 200 {"url":"...","format":"mp3","source":"generated","expires_in_seconds":300}The returned link expires in 300 seconds, so download or play it right away and keep your own copy of lines you replay. Keep the key on a server: web, mobile and kiosk setups.
Price, storage and privacy
| Price | $6 per 1M characters (list); your first 100,000 characters across all voices are free |
| What is stored | A sentence is stored privately only after it has been requested twice; the first request is generated and returned without being kept |
| How long | Deleted after 75 days without use; each reuse restarts the 75 days |
| Where it is made | By our voice provider (Chatterbox on DeepInfra); not available to other accounts |
| Counted how | Characters, billed on every request, repeats included. See billing explained |
Delivery styles (preview, not in production)
Happy, excited, sorry, calm, serious or neutral exist as a private preview in the development environment only. They are English only, would cost $17 per million characters, and are not enabled in production. We will not list them as available until they are.
Consent and responsible use
The portal asks you to confirm you have the right to use the voice you submit. Do not clone someone who has not agreed. Where people could assume they are talking to a person, say that the voice is automated. We do not build lines that pretend an automated voice is a specific human.
Next: how generation works · billing explained · dashboard tour