One endpoint to speak a phrase. Machine-readable files are provided for AI assistants.
Create a key in the dashboard, then request audio.
curl https://speakvora.com/api/v1/speak \
-H "Authorization: Bearer $SPEAKVORA_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Aim assist is off.","voice":"af_heart","lang":"en"}' \
POST /v1/speak returns JSON with a url to the audio (m4a). Full schema in the OpenAPI spec.
Text that is not in the phrase library is voiced on demand with the Kokoro voices (English, Spanish, French, Portuguese, Hindi, Japanese and Mandarin; non-English results are marked native_review_pending) and returned as an mp3 url. A brand-new sentence takes seconds (6 to 9 seconds in two observations on 8 October 2026); the same text afterwards returns quickly. Premium voices use the phrase library only.
Pass language as an ISO code or auto. Ambiguous text is rejected rather than guessed. Non-English clips carry native_review_pending until a native speaker approves them.
Every account starts with 100,000 free characters, with no card. The portal shows how much is left. When it is used up, requests return 402 free_allowance_used until you add a payment method in the portal. Library phrases count by their length like any other text.
The home page and playground speak through POST /api/demo/speak: up to 140 characters, nine voices, and a small hourly limit per visitor. Use it to hear a voice. Use your key for real work.
Speakvora runs a Model Context Protocol server, so an AI assistant can speak text, list voices and search phrases without any code from you. Endpoint: https://speakvora.com/api/mcp (streamable HTTP). Send your key as x-api-key or Authorization: Bearer.
# Claude Code
claude mcp add --transport http speakvora https://speakvora.com/api/mcp --header "x-api-key: YOUR_KEY"
# Cursor and other clients that read an mcp.json file
{ "mcpServers": { "speakvora": { "url": "https://speakvora.com/api/mcp", "headers": { "x-api-key": "YOUR_KEY" } } } }
Tools: speak (text, voice, lang) returns an audio URL; list_voices; search_phrases (query, lang, limit). Client setup differs between tools, so check yours for how it stores HTTP servers and headers.
GET /api/v1/search?q=refund+order&lang=en&limit=10 returns phrases that contain your words, ranked, with the voices that have them. A library phrase plays instantly and costs nothing to generate. Search is word-based today; meaning-based search is planned.
GET /api/v1/packs lists downloadable zips. Each holds the audio and a manifest.json that maps sentence to file. Browse them on the packs page.
Customers can have a voice of their own, made in the portal in three ways: record a paragraph in the browser, upload a recording (10 to 60 seconds of one clear voice), or describe and tune a voice with meters (gender, age, pitch, pace, warmth, energy) and optional notes. A designed voice is first rendered as a reference sample, then saved like a recording, so it sounds the same every time. Designing a voice to imitate a named real person is refused. A recorded or uploaded voice needs a confirmation that it is yours or you have the speaker's permission.
A voice becomes ready only after it passes a check: we speak a test sentence with it and listen to the result. Then about two dozen common support lines (greetings, thanks, clarification, waiting, closing) are prepared in the background, with progress shown, so those play instantly. Everything else is generated on first use and reused. Each account can keep 3 voices. You can preview, replace the sample, or delete a voice; deleting also removes every private recording made with it.
| Language with a custom voice | Status (measured with a cloned voice) |
|---|---|
| English | Supported |
| French, German | Beta: understandable, with small errors |
| Spanish, Hindi, Chinese, Japanese | Not offered: not intelligible in testing. Use a catalog voice. |
Custom voices are generated by our voice provider (Chatterbox on DeepInfra). They are not available to other accounts. Spoken laughter ([laugh], [chuckle], [cough]) exists only for English custom voices and is off unless you turn it on. Personality does not add effects.
Status: private preview. Styles are switched on in the development environment and are not enabled in production yet. Any ready custom voice can speak with a style: happy, excited, sorry, calm, serious or neutral (default). Styles are English only and cost $17 per million characters (plain custom voice: $6). Pick the style in the portal, or send it in the request:
POST /api/v1/speak header x-api-key: <key>
{ "text": "I am so sorry about the mix up with your order.", "voice": "cv_07b92a5f8b21", "style": "sorry" }
200 { "url": "https://...", "format": "wav", "source": "generated|cache", "style": "sorry", "style_applied": true, "engine": "mimo", "expires_in_seconds": 300 }
Styled audio is a WAV file at a signed link valid for 5 minutes. A brand-new styled sentence takes about 4 to 10 seconds; a repeat is instant. If our primary styling engine is unavailable, a backup engine keeps your voice speaking but cannot apply the style: the reply then says style_applied: false and the sentence is billed at the plain $6 rate. Accents cannot be requested: they did not sound reliable in our tests.
Caching and retention. A custom-voice sentence is stored only once it has been asked for twice; the first request is generated and returned without being kept. Stored recordings are private to your account and deleted after 75 days without use (each reuse restarts the 75 days).
Send audio, get text. Billed per second of audio at $0.0015 per minute; your first 60 minutes are free.
POST /api/v1/listen header x-api-key: <key>
{ "audio_base64": "<base64>", "format": "wav|mp3|m4a|ogg|flac|webm", "language": "en" } # language is optional
200 { "text": "Your refund will arrive in three to five business days.", "seconds": 3, "language": "en" }Audio must be 1 KB to 4.5 MB per request; send longer recordings in pieces. Not offered yet: speaker labels and word timestamps. Errors: 400 invalid_request, 402 free_allowance_used, 502 upstream_unavailable.
For live transcription, streamed speech and streamed agent turns, open a WebSocket and send JSON text frames. The first message must authenticate.
wss://48f8pzwtee.execute-api.us-east-1.amazonaws.com/v1 (the address of this environment)
> {"type":"auth","key":"YOUR_KEY"} < {"type":"ready"}
Live transcription
> {"type":"listen.start","language":"en","rate":16000,"interim":false} # optional: endpoint_ms (700), vad_threshold (400)
> {"type":"audio","seq":0,"data":"<base64 PCM16 mono, 100-500 ms>"} # seq counts up from 0
< {"type":"final","text":"...","seconds":4,"latency_ms":830} # about 0.8 s after speech stops
> {"type":"listen.end"} # flush
Streamed speech
> {"type":"speak","id":"1","text":"Your order has shipped. It arrives Friday.","voice":"af_heart","lang":"en"}
< {"type":"audio","index":0,"text":"Your order has shipped.","url":"https://...","source":"live"} # one per sentence, in order
< {"type":"done","sentences":2,"characters":50}
Agent turn
> {"type":"agent.turn","agent":"ag_...","text":"Can I move my appointment?","session":"s1"}End-of-speech detection uses loudness: raise vad_threshold in noisy rooms. With interim: true you also receive partial results about every 2 seconds of speech; these re-transcribe the buffer and are billed as audio seconds. Audio comes back as links to play, not as raw streamed audio.
| Product | Price |
|---|---|
| Catalog voices (Kokoro) | $0.70 per 1M characters |
| Your own voice | $6 per 1M characters |
| Styled voice (happy, excited, sorry, calm, serious) | $17 per 1M characters |
| Speech to text | $0.0015 per minute |
| Voice agents | $10 per 1,000 turns |
| Offline phrase packs | US$24 to US$399 one time per pack |
Free: 100,000 characters, 60 minutes of speech to text and 200 agent turns per account. Characters are counted on the text you send. Prices are list prices in US dollars.
A pack is a zip of ready-made audio for one topic, language and voice, to ship inside your app so common lines play with no network and no per-play cost. Each pack holds manifest.json (every sentence with its audio file), phrases.csv, the audio (.m4a) and a licence. Packs are a one-time purchase (US$24 to US$399 by size and topic), rebuilt as the library grows, and buyers can download newer versions.
GET /api/v1/packs public list: id, title, phrases, size, price_usd, sample_url, examples
GET /packs/{id}/sample.zip free sample, 25 phrases, no key
GET /api/v1/packs/{id}/download x-api-key 200 {"url": "...zip (valid 5 minutes)"} once the pack is bought on the account, else 402 pack_not_purchased
Buy a pack on the packs page; it is paid by card and appears under Plan and billing in the portal.
Open Agent builder in the portal and describe the agent in your own words, for example: a warm receptionist who answers briefly, helps with appointments, and speaks in my voice. The chat fills in an editable configuration next to it: name, purpose, instructions, greeting, personality (professional, friendly, calm, energetic, empathetic or your own wording), language, voice, response length, speaking style, business knowledge and FAQs, actions (tools), memory and interruption. Business facts such as prices and hours come only from the knowledge you provide; the agent says it does not know otherwise. Personality changes the wording. Personality changes wording only; to change how a voice sounds, use a delivery style (see Your own voice).
Press Generate API to publish. Each publish saves a numbered, unchangeable version at the same endpoint, and the first publish gives you a key that works only for that agent (shown once; rotate it any time). Customers never need a DeepInfra account.
POST /api/v1/agents/{agent_id}/turn header x-api-key: <agent key> (or Authorization: Bearer <session token>)
{ "session": "call-1", "text": "Can I move my appointment to Thursday?" }
or { "session": "call-1", "audio_b64": "<WAV, base64>" } # we transcribe it first
or { "session": "call-1", "start": true } # speaks the greeting
200 { "session": "call-1", "generation": "9f2c...", "reply": "Of course. What day works for you?",
"sentences": [ { "index": 0, "text": "Of course.", "audio_url": "https://...", "source": "library|cache|generated|live" } ],
"timings": { "stt_ms": 0, "llm_ms": 410, "first_audio_ms": 1650, "total_ms": 2840 }, "agent": "ag_...", "version": 3 }
Browser apps: keep the agent key on your server and call POST /api/v1/agents/{agent_id}/sessions to get a token valid for 10 minutes; the browser sends it as Authorization: Bearer st..... A token (or an agent key) can only reach that one agent.
Interruption: play sentences[].audio_url in order. If the caller talks over the agent, stop playback and send the next turn with "interrupted": {"generation": "...", "played": 1}, or call POST /api/v1/agents/{agent_id}/interrupt with {session, generation, played}. Sentences not yet produced are discarded and the agent's memory keeps only what was actually played.
What this is and is not: each turn streams the model's reply internally and speaks it sentence by sentence, but the HTTP response arrives after the sentences are produced, so the first audio link is available when the whole turn returns (typically 2 to 4 seconds with a custom voice; catalog-voice sentences that are not in the library can take longer). Live token streaming over a WebSocket is not available yet. Transcription is a file upload, not a continuous stream, so this is a fast turn-based conversation, not a full-duplex phone call. There are no phone numbers. First 200 turns free; then $10 per 1,000 turns in early access.
Live checks and daily uptime are on the status page. GET /api/status returns the same data.
Official clients for JavaScript (Node 18+ and browsers) and Python (3.8+), with no dependencies. They cover speech, prepare, speech to text, voices, search, offline packs, voice agents, live streaming and webhook checks, and return errors with a plain-language fix.
# JavaScript # Python
npm install https://speakvora.com/sdk/speakvora-0.2.0.tgz pip install https://speakvora.com/sdk/speakvora-0.2.0.tar.gz
const { Speakvora } = require("speakvora"); from speakvora import Speakvora
const sv = new Speakvora(process.env.SPEAKVORA_KEY); sv = Speakvora(os.environ["SPEAKVORA_KEY"])
const clip = await sv.speak("Your order has shipped."); clip = sv.speak("Your order has shipped.")
await sv.prepare(["Your table is ready."]); sv.prepare(["Your table is ready."])
const { text } = await sv.listen(wavBuffer); text = sv.listen(wav_bytes)["text"]
const turn = await sv.agent(id, agentKey).turn({ session: "c1", text: "Hours?" })
Single-file copies for pasting into a project: speakvora.js and speakvora.py. Keep your key on a server; for browsers mint a short-lived token with agent(id).sessions(). The npm and PyPI listings will follow the production launch.
Add a URL in the portal and we POST signed JSON events: test.ping, key.created, key.revoked. Each request carries x-speakvora-timestamp and x-speakvora-signature, the hex HMAC-SHA256 of timestamp + "." + raw body using your signing secret. Use Speakvora.verifyWebhook or verify_webhook from the SDKs, and reject anything older than five minutes.
Use GET /api/v1/voices and GET /api/v1/languages for the live list. Voices that can speak any text live:
| Language | lang | Voices |
|---|---|---|
| English | en | af_heart, af_bella, af_nicole, af_sarah, af_sky, am_adam, am_michael |
| Spanish | es | ef_dora, em_alex |
| French | fr | ff_siwis |
| Portuguese | pt | pf_dora, pm_alex |
| Hindi | hi | hf_alpha, hm_omega |
| Japanese | ja | jf_alpha, jm_kumo |
| Mandarin Chinese | zh | zf_001, zm_009 |
Other voices and languages play phrases from the library only, until live generation reaches them. The language you send must match the voice.
429 rate_limited with retry_after_seconds.live_rate_limited. Speed: library phrases and repeats are instant. A brand-new catalog sentence takes about 3 to 5 seconds, and the first one after a quiet period can take up to 30 seconds. A new custom-voice sentence takes about 2 seconds (styled: 4 to 10 seconds); a repeat is a cache hit (about 0.3 s).Every error reply has error, message, fix and docs fields. The status codes are the standard ones.
| Status | error | Meaning and fix |
|---|---|---|
| 401 | unauthorized | Key missing, wrong or revoked. Send it in x-api-key. Check for spaces or quotes around it. |
| 400 | invalid_json | Body is not valid JSON. Send content-type: application/json and check commas and quotes. |
| 400 | text_required_max_500_chars | Text empty or over 500 characters. Split it. |
| 400 | voice_required | Add a voice id. |
| 422 | custom_voice_error | The custom voice could not be used. Read detail: styles are English only, the voice must be ready, the id must be yours. |
| 422 | style_not_available | Styles are not switched on for this environment. |
| 400 | invalid_request | Speech-to-text or streaming request malformed. Read detail. |
| 404 | voice_not_live | That voice cannot speak new text. Pick one from the table above. |
| 422 | language_not_supported_for_voice | lang does not match the voice. The reply says which language the voice speaks. |
| 422 | language_not_supported | Language not available yet. The reply lists the ones that are. |
| 422 | invalid_text | Text rejected. Read reasons_explained: usually digits in a non-English language, a placeholder like (name), or a safety claim. |
| 404 | voice_not_available_for_text | The phrase exists, but not in that voice. The reply lists the voices that have it. |
| 404 | not_in_library | Not a library phrase and live generation could not run for it. |
| 422 | generation_failed_quality | The audio was too short or broken, so it was not served. Add a full stop or reword. |
| 429 | (Too Many Requests) | Over the rate limit. Wait a moment and retry with a short backoff. |
| 502/503/504 | (gateway timeout) | A new sentence took longer than the time limit, usually the first request after a quiet period. Retry the same request. |
x-api-key, with no Bearer unless you use Authorization.lang: "en" fails.url directly. Browsers need a click or tap before they will autoplay audio.Why was a key rejected right after I created it? Check for a trailing space or newline when you pasted it. Keys start with sv_.
Can I send personal data? No. Use placeholders on your side and speak the personal part another way.
Are non-English voices final? They are labelled native_review_pending until a native speaker approves the language.
Two different things happen to audio, and the rules differ:
Catalog voices (/v1/speak, phrase library) | Custom voices and agents | |
|---|---|---|
| Who can fetch the audio | Anyone with the link. Links are long and unguessable but public, and the same text in the same voice returns the same file. | Only your account. Links are signed and last 5 minutes. |
| How long it is kept | Library clips: served with a one-year cache lifetime. Live-generated clips: kept for reuse. | 75 days after the last time it was used; each reuse restarts the 75 days. Expired audio is refused and deleted. |
| Personal data in the text | Do not send it. Send generic sentences and speak the personal part another way. | Allowed if you have the right to use it. It stays private to your account. |
Other things to know: conversation memory for agents is kept for 24 hours. Your text, and recordings you submit to make a voice, are processed by our AI provider (DeepInfra), which lists zero retention on the models we use; we do not keep your voice sample after the voice is made. You can delete one voice, one recording, or all private audio from the portal. Private audio has a storage quota (500 MB by default). Text with unresolved placeholders such as (name) or {name} is rejected, so build the real sentence on your side. Whether you may record, clone or process someone's voice or personal data is your responsibility under the laws that apply to you; we are not a law firm and this is not legal advice.
Point your coding assistant at /docs/llms.txt and the OpenAPI spec; they contain everything needed to wire up the API.