Use your own voice in your apps
Make the voice once in the portal, then speak with it from a web app, a mobile app or a kiosk without exposing your key.
Record your voice, name it and use it everywhere with the same API call. Three setups, one rule: the key stays on a server you control.
1. Make the voice (once)
- Open the portal and choose Make your own voice.
- Record a paragraph in the browser, upload 10 to 60 seconds of one clear voice, or describe a voice with the meters (gender, age, pitch, pace, warmth, energy).
- Confirm you have the right to use the voice, and give it a name.
- Wait for ready. The portal speaks a test sentence and listens to it with speech recognition before it marks the voice ready, then prepares about two dozen common lines so they play fast.
- Copy the voice id; it starts with
cv_.
2. Speak with it from your backend
cURL
curl -X POST https://speakvora.com/api/v1/speak -H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
-d '{"text":"Thanks for calling.","voice":"cv_YOUR_VOICE_ID","lang":"en"}'
# 200 {"url":"...","format":"mp3","source":"generated","expires_in_seconds":300,"characters":19,"voice":"cv_...","lang":"en"}3. Pick your setup
Web app
Browser calls your /api/say route; your server calls Speakvora and returns the link; the browser plays it. See add speech to a web app, then change voice to your cv_ id.
Mobile app
Never embed the key in the app. Your backend returns the link; the app downloads and plays it. Download within 300 seconds, because the link expires.
Kiosk or device
Generate your fixed lines at build time on a server, store the files with the device image, and call the API only for new sentences. See store and reuse audio.
What to expect
- Languages: English supported; French and German beta; Spanish, Hindi, Chinese and Japanese are not offered for custom voices (not intelligible in our testing).
- Speed: a new sentence took about 2.3 seconds in an earlier test and a repeat about 0.3 seconds. Those are single observations, not guarantees.
- Storage: a sentence is stored privately only after it has been requested twice; stored recordings are deleted after 75 days without use, and each reuse restarts the 75 days.
- Limits: 3 voices per account; replacing the sample changes the voice version; deleting a voice also deletes every private recording made with it.
- Price: $6 per million characters (list).
- Responsible use: only clone a voice you have the right to use; say that a voice is automated where people could assume otherwise.
The full picture: your own voice, in depth.