Use your own voice in your apps

Record yourself, name the voice, and speak with it through the same API call you use for everything else. English supported today.

Make it in three steps

1. Give us a voice

Record a paragraph in the browser, upload 10 to 60 seconds of one clear voice, or describe a voice with meters for gender, age, pitch, pace, warmth and energy.

2. Name it

Confirm you have the right to use the voice. It is private to your account and never available to other accounts.

3. Speak

Send its cv_ id as the voice in /v1/speak. That is the whole integration.

What happens behind the scenes

  • Verification before use. A voice is marked ready only after we speak a test sentence with it, check it is real audio, and check with speech recognition that the words were heard. A voice that fails is removed.
  • Starter lines. About two dozen common lines (greetings, thanks, holds, closings) are prepared in the background so they play fast.
  • Designed voices are first rendered as a reference sample, then saved like a recording, so they sound the same every time.
  • Limits: 3 voices per account. Replacing the sample bumps the voice version. Deleting a voice removes it at our provider and deletes every private recording made with it.

Languages (measured, not promised)

LanguageStatus
EnglishSupported
French, GermanBeta: understandable, with small errors
Spanish, Hindi, Chinese, JapaneseNot offered: not intelligible in our testing. Use a catalog voice

Use it in your code

cURL

curl -X POST https://speakvora.com/api/v1/speak -H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
  -d '{"text":"Thanks for calling.","voice":"cv_YOUR_VOICE_ID","lang":"en"}'

# 200 {"url":"...","format":"mp3","source":"generated","expires_in_seconds":300}

The returned link expires in 300 seconds, so download or play it right away and keep your own copy of lines you replay. Keep the key on a server: web, mobile and kiosk setups.

Price, storage and privacy

Price$6 per 1M characters (list); your first 100,000 characters across all voices are free
What is storedA sentence is stored privately only after it has been requested twice; the first request is generated and returned without being kept
How longDeleted after 75 days without use; each reuse restarts the 75 days
Where it is madeBy our voice provider (Chatterbox on DeepInfra); not available to other accounts
Counted howCharacters, billed on every request, repeats included. See billing explained

Delivery styles (preview, not in production)

Happy, excited, sorry, calm, serious or neutral exist as a private preview in the development environment only. They are English only, would cost $17 per million characters, and are not enabled in production. We will not list them as available until they are.

Consent and responsible use

The portal asks you to confirm you have the right to use the voice you submit. Do not clone someone who has not agreed. Where people could assume they are talking to a person, say that the voice is automated. We do not build lines that pretend an automated voice is a specific human.

Next: how generation works · billing explained · dashboard tour