Learn: answers about Speakvora
Short, sourced answers about the Speakvora speech API: how-tos, use cases and common questions. Every page is checked against our product facts before it is published.
What is a text-to-speech API?
What is a text-to-speech API?
What is the difference between text to speech and speech to text?
Text to speech vs speech to text
How is text-to-speech priced?
How text-to-speech pricing works
What is an MCP server for voice?
What is an MCP server for voice?
Custom Voice Creation: Record, Upload, or Design
Create custom voices in Speakvora: record a paragraph, upload speech, or design with tuning meters. English supported; French and German beta. Max 3 voices per account.
Create, scope and revoke API keys safely
Manage up to 5 API keys per account with agent scoping, server-side storage, and webhook notifications for key creation and revocation.
WebSocket Streaming: Audio Links and Live Transcription
Speakvora streams speech and transcription over WebSocket with audio returned as URLs per sentence, not raw bytes. Live transcription and agent turns supported.
Download an industry pack after purchase
Get a signed download link for your purchased Speakvora pack using the API, then extract the audio, manifest and phrases.
Audio and Voice Recording Retention
How Speakvora stores, secures and deletes your custom voice recordings, agent audio, and conversation memory.
API Keys and Scoped Access
Manage up to 5 API keys per account. Scoped keys restrict access to a single agent. Learn authentication, key limits, and webhook events.
How Speakvora Text-to-Speech Billing Works
Speakvora charges per character for text-to-speech: $0.70 per 1M for catalog voices, $6 per 1M for custom voices. Library hits are billed every time.
SDKs and client libraries for Speakvora
Speakvora offers single-file JavaScript and Python helpers. Call the API directly from any language using HTTP.
Play Speakvora speech in a browser safely
Call /v1/speak from your backend and pass the audio URL to the browser. Agent tokens work in the browser directly but are limited to agents.
Prepare phrases ahead of time with /v1/prepare
Use POST /v1/prepare to queue up to 200 texts for background generation without billing. Ready phrases return instantly; queued texts are ready in a few seconds.
Turn text into speech with one curl request
Make a single POST request to /v1/speak with text and a voice ID to get a URL to audio. Authenticate with x-api-key header.
Write numbers as words for non-English speech
English accepts digits in text-to-speech, but Spanish, French, Portuguese, Hindi, Japanese and Chinese require numbers written as words to generate speech.