Use your own voice in your apps

Make the voice once in the portal, then speak with it from a web app, a mobile app or a kiosk without exposing your key.

Record your voice, name it and use it everywhere with the same API call. Three setups, one rule: the key stays on a server you control.

1. Make the voice (once)

  1. Open the portal and choose Make your own voice.
  2. Record a paragraph in the browser, upload 10 to 60 seconds of one clear voice, or describe a voice with the meters (gender, age, pitch, pace, warmth, energy).
  3. Confirm you have the right to use the voice, and give it a name.
  4. Wait for ready. The portal speaks a test sentence and listens to it with speech recognition before it marks the voice ready, then prepares about two dozen common lines so they play fast.
  5. Copy the voice id; it starts with cv_.

2. Speak with it from your backend

cURL

curl -X POST https://speakvora.com/api/v1/speak -H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
  -d '{"text":"Thanks for calling.","voice":"cv_YOUR_VOICE_ID","lang":"en"}'

# 200 {"url":"...","format":"mp3","source":"generated","expires_in_seconds":300,"characters":19,"voice":"cv_...","lang":"en"}

3. Pick your setup

Web app

Browser calls your /api/say route; your server calls Speakvora and returns the link; the browser plays it. See add speech to a web app, then change voice to your cv_ id.

Mobile app

Never embed the key in the app. Your backend returns the link; the app downloads and plays it. Download within 300 seconds, because the link expires.

Kiosk or device

Generate your fixed lines at build time on a server, store the files with the device image, and call the API only for new sentences. See store and reuse audio.

What to expect

  • Languages: English supported; French and German beta; Spanish, Hindi, Chinese and Japanese are not offered for custom voices (not intelligible in our testing).
  • Speed: a new sentence took about 2.3 seconds in an earlier test and a repeat about 0.3 seconds. Those are single observations, not guarantees.
  • Storage: a sentence is stored privately only after it has been requested twice; stored recordings are deleted after 75 days without use, and each reuse restarts the 75 days.
  • Limits: 3 voices per account; replacing the sample changes the voice version; deleting a voice also deletes every private recording made with it.
  • Price: $6 per million characters (list).
  • Responsible use: only clone a voice you have the right to use; say that a voice is automated where people could assume otherwise.

The full picture: your own voice, in depth.