Learn: answers about Speakvora

Short, sourced answers about the Speakvora speech API: how-tos, use cases and common questions. Every page is checked against our product facts before it is published.

What is a text-to-speech API?

What is a text-to-speech API?

What is the difference between text to speech and speech to text?

Text to speech vs speech to text

How is text-to-speech priced?

How text-to-speech pricing works

What is an MCP server for voice?

What is an MCP server for voice?

Custom Voice Creation: Record, Upload, or Design

Create custom voices in Speakvora: record a paragraph, upload speech, or design with tuning meters. English supported; French and German beta. Max 3 voices per account.

Create, scope and revoke API keys safely

Manage up to 5 API keys per account with agent scoping, server-side storage, and webhook notifications for key creation and revocation.

WebSocket Streaming: Audio Links and Live Transcription

Speakvora streams speech and transcription over WebSocket with audio returned as URLs per sentence, not raw bytes. Live transcription and agent turns supported.

Download an industry pack after purchase

Get a signed download link for your purchased Speakvora pack using the API, then extract the audio, manifest and phrases.

Audio and Voice Recording Retention

How Speakvora stores, secures and deletes your custom voice recordings, agent audio, and conversation memory.

API Keys and Scoped Access

Manage up to 5 API keys per account. Scoped keys restrict access to a single agent. Learn authentication, key limits, and webhook events.

How Speakvora Text-to-Speech Billing Works

Speakvora charges per character for text-to-speech: $0.70 per 1M for catalog voices, $6 per 1M for custom voices. Library hits are billed every time.

SDKs and client libraries for Speakvora

Speakvora offers single-file JavaScript and Python helpers. Call the API directly from any language using HTTP.

Play Speakvora speech in a browser safely

Call /v1/speak from your backend and pass the audio URL to the browser. Agent tokens work in the browser directly but are limited to agents.

Prepare phrases ahead of time with /v1/prepare

Use POST /v1/prepare to queue up to 200 texts for background generation without billing. Ready phrases return instantly; queued texts are ready in a few seconds.

Turn text into speech with one curl request

Make a single POST request to /v1/speak with text and a voice ID to get a URL to audio. Authenticate with x-api-key header.

Write numbers as words for non-English speech

English accepts digits in text-to-speech, but Spanish, French, Portuguese, Hindi, Japanese and Chinese require numbers written as words to generate speech.