Developer docs

One endpoint to speak a phrase. Machine-readable files are provided for AI assistants.

Quick start

Create a key in the dashboard, then request audio.

curl https://speakvora.com/api/v1/speak \
  -H "Authorization: Bearer $SPEAKVORA_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"Aim assist is off.","voice":"af_heart","lang":"en"}' \
  

Speak

POST /v1/speak returns JSON with a url to the audio (m4a). Full schema in the OpenAPI spec.

Any text, live

Text that is not in the phrase library is voiced on demand with the Kokoro voices (English, Spanish, French, Portuguese, Hindi, Japanese and Mandarin; non-English results are marked native_review_pending) and returned as an mp3 url. A brand-new sentence takes seconds (6 to 9 seconds in two observations on 8 October 2026); the same text afterwards returns quickly. Premium voices use the phrase library only.

Languages

Pass language as an ISO code or auto. Ambiguous text is rejected rather than guessed. Non-English clips carry native_review_pending until a native speaker approves them.

Free allowance

Every account starts with 100,000 free characters, with no card. The portal shows how much is left. When it is used up, requests return 402 free_allowance_used until you add a payment method in the portal. Library phrases count by their length like any other text.

Try it without an account

The home page and playground speak through POST /api/demo/speak: up to 140 characters, nine voices, and a small hourly limit per visitor. Use it to hear a voice. Use your key for real work.

MCP: use Speakvora from Claude, Cursor and other AI tools

Speakvora runs a Model Context Protocol server, so an AI assistant can speak text, list voices and search phrases without any code from you. Endpoint: https://speakvora.com/api/mcp (streamable HTTP). Send your key as x-api-key or Authorization: Bearer.

# Claude Code
claude mcp add --transport http speakvora https://speakvora.com/api/mcp --header "x-api-key: YOUR_KEY"

# Cursor and other clients that read an mcp.json file
{ "mcpServers": { "speakvora": { "url": "https://speakvora.com/api/mcp", "headers": { "x-api-key": "YOUR_KEY" } } } }

Tools: speak (text, voice, lang) returns an audio URL; list_voices; search_phrases (query, lang, limit). Client setup differs between tools, so check yours for how it stores HTTP servers and headers.

GET /api/v1/search?q=refund+order&lang=en&limit=10 returns phrases that contain your words, ranked, with the voices that have them. A library phrase plays instantly and costs nothing to generate. Search is word-based today; meaning-based search is planned.

Offline packs

GET /api/v1/packs lists downloadable zips. Each holds the audio and a manifest.json that maps sentence to file. Browse them on the packs page.

Your own voice

Customers can have a voice of their own, made in the portal in three ways: record a paragraph in the browser, upload a recording (10 to 60 seconds of one clear voice), or describe and tune a voice with meters (gender, age, pitch, pace, warmth, energy) and optional notes. A designed voice is first rendered as a reference sample, then saved like a recording, so it sounds the same every time. Designing a voice to imitate a named real person is refused. A recorded or uploaded voice needs a confirmation that it is yours or you have the speaker's permission.

A voice becomes ready only after it passes a check: we speak a test sentence with it and listen to the result. Then about two dozen common support lines (greetings, thanks, clarification, waiting, closing) are prepared in the background, with progress shown, so those play instantly. Everything else is generated on first use and reused. Each account can keep 3 voices. You can preview, replace the sample, or delete a voice; deleting also removes every private recording made with it.

Language with a custom voiceStatus (measured with a cloned voice)
EnglishSupported
French, GermanBeta: understandable, with small errors
Spanish, Hindi, Chinese, JapaneseNot offered: not intelligible in testing. Use a catalog voice.

Custom voices are generated by our voice provider (Chatterbox on DeepInfra). They are not available to other accounts. Spoken laughter ([laugh], [chuckle], [cough]) exists only for English custom voices and is off unless you turn it on. Personality does not add effects.

Delivery styles

Status: private preview. Styles are switched on in the development environment and are not enabled in production yet. Any ready custom voice can speak with a style: happy, excited, sorry, calm, serious or neutral (default). Styles are English only and cost $17 per million characters (plain custom voice: $6). Pick the style in the portal, or send it in the request:

POST /api/v1/speak   header x-api-key: <key>
{ "text": "I am so sorry about the mix up with your order.", "voice": "cv_07b92a5f8b21", "style": "sorry" }

200 { "url": "https://...", "format": "wav", "source": "generated|cache", "style": "sorry", "style_applied": true, "engine": "mimo", "expires_in_seconds": 300 }

Styled audio is a WAV file at a signed link valid for 5 minutes. A brand-new styled sentence takes about 4 to 10 seconds; a repeat is instant. If our primary styling engine is unavailable, a backup engine keeps your voice speaking but cannot apply the style: the reply then says style_applied: false and the sentence is billed at the plain $6 rate. Accents cannot be requested: they did not sound reliable in our tests.

Caching and retention. A custom-voice sentence is stored only once it has been asked for twice; the first request is generated and returned without being kept. Stored recordings are private to your account and deleted after 75 days without use (each reuse restarts the 75 days).

Speech to text

Send audio, get text. Billed per second of audio at $0.0015 per minute; your first 60 minutes are free.

POST /api/v1/listen   header x-api-key: <key>
{ "audio_base64": "<base64>", "format": "wav|mp3|m4a|ogg|flac|webm", "language": "en" }    # language is optional

200 { "text": "Your refund will arrive in three to five business days.", "seconds": 3, "language": "en" }

Audio must be 1 KB to 4.5 MB per request; send longer recordings in pieces. Not offered yet: speaker labels and word timestamps. Errors: 400 invalid_request, 402 free_allowance_used, 502 upstream_unavailable.

Live streaming (WebSocket)

For live transcription, streamed speech and streamed agent turns, open a WebSocket and send JSON text frames. The first message must authenticate.

wss://48f8pzwtee.execute-api.us-east-1.amazonaws.com/v1      (the address of this environment)

> {"type":"auth","key":"YOUR_KEY"}                                   < {"type":"ready"}

Live transcription
> {"type":"listen.start","language":"en","rate":16000,"interim":false}   # optional: endpoint_ms (700), vad_threshold (400)
> {"type":"audio","seq":0,"data":"<base64 PCM16 mono, 100-500 ms>"}      # seq counts up from 0
< {"type":"final","text":"...","seconds":4,"latency_ms":830}            # about 0.8 s after speech stops
> {"type":"listen.end"}                                               # flush

Streamed speech
> {"type":"speak","id":"1","text":"Your order has shipped. It arrives Friday.","voice":"af_heart","lang":"en"}
< {"type":"audio","index":0,"text":"Your order has shipped.","url":"https://...","source":"live"}   # one per sentence, in order
< {"type":"done","sentences":2,"characters":50}

Agent turn
> {"type":"agent.turn","agent":"ag_...","text":"Can I move my appointment?","session":"s1"}

End-of-speech detection uses loudness: raise vad_threshold in noisy rooms. With interim: true you also receive partial results about every 2 seconds of speech; these re-transcribe the buffer and are billed as audio seconds. Audio comes back as links to play, not as raw streamed audio.

Prices

ProductPrice
Catalog voices (Kokoro)$0.70 per 1M characters
Your own voice$6 per 1M characters
Styled voice (happy, excited, sorry, calm, serious)$17 per 1M characters
Speech to text$0.0015 per minute
Voice agents$10 per 1,000 turns
Offline phrase packsUS$24 to US$399 one time per pack

Free: 100,000 characters, 60 minutes of speech to text and 200 agent turns per account. Characters are counted on the text you send. Prices are list prices in US dollars.

Offline phrase packs

A pack is a zip of ready-made audio for one topic, language and voice, to ship inside your app so common lines play with no network and no per-play cost. Each pack holds manifest.json (every sentence with its audio file), phrases.csv, the audio (.m4a) and a licence. Packs are a one-time purchase (US$24 to US$399 by size and topic), rebuilt as the library grows, and buyers can download newer versions.

GET /api/v1/packs                                    public list: id, title, phrases, size, price_usd, sample_url, examples
GET /packs/{id}/sample.zip                           free sample, 25 phrases, no key
GET /api/v1/packs/{id}/download   x-api-key          200 {"url": "...zip (valid 5 minutes)"} once the pack is bought on the account, else 402 pack_not_purchased

Buy a pack on the packs page; it is paid by card and appears under Plan and billing in the portal.

Voice agents

Open Agent builder in the portal and describe the agent in your own words, for example: a warm receptionist who answers briefly, helps with appointments, and speaks in my voice. The chat fills in an editable configuration next to it: name, purpose, instructions, greeting, personality (professional, friendly, calm, energetic, empathetic or your own wording), language, voice, response length, speaking style, business knowledge and FAQs, actions (tools), memory and interruption. Business facts such as prices and hours come only from the knowledge you provide; the agent says it does not know otherwise. Personality changes the wording. Personality changes wording only; to change how a voice sounds, use a delivery style (see Your own voice).

Press Generate API to publish. Each publish saves a numbered, unchangeable version at the same endpoint, and the first publish gives you a key that works only for that agent (shown once; rotate it any time). Customers never need a DeepInfra account.

POST /api/v1/agents/{agent_id}/turn        header x-api-key: <agent key>  (or Authorization: Bearer <session token>)
{ "session": "call-1", "text": "Can I move my appointment to Thursday?" }
        or  { "session": "call-1", "audio_b64": "<WAV, base64>" }      # we transcribe it first
        or  { "session": "call-1", "start": true }                     # speaks the greeting

200 { "session": "call-1", "generation": "9f2c...", "reply": "Of course. What day works for you?",
      "sentences": [ { "index": 0, "text": "Of course.", "audio_url": "https://...", "source": "library|cache|generated|live" } ],
      "timings": { "stt_ms": 0, "llm_ms": 410, "first_audio_ms": 1650, "total_ms": 2840 }, "agent": "ag_...", "version": 3 }

Browser apps: keep the agent key on your server and call POST /api/v1/agents/{agent_id}/sessions to get a token valid for 10 minutes; the browser sends it as Authorization: Bearer st..... A token (or an agent key) can only reach that one agent.

Interruption: play sentences[].audio_url in order. If the caller talks over the agent, stop playback and send the next turn with "interrupted": {"generation": "...", "played": 1}, or call POST /api/v1/agents/{agent_id}/interrupt with {session, generation, played}. Sentences not yet produced are discarded and the agent's memory keeps only what was actually played.

What this is and is not: each turn streams the model's reply internally and speaks it sentence by sentence, but the HTTP response arrives after the sentences are produced, so the first audio link is available when the whole turn returns (typically 2 to 4 seconds with a custom voice; catalog-voice sentences that are not in the library can take longer). Live token streaming over a WebSocket is not available yet. Transcription is a file upload, not a continuous stream, so this is a fast turn-based conversation, not a full-duplex phone call. There are no phone numbers. First 200 turns free; then $10 per 1,000 turns in early access.

Status

Live checks and daily uptime are on the status page. GET /api/status returns the same data.

SDKs

Official clients for JavaScript (Node 18+ and browsers) and Python (3.8+), with no dependencies. They cover speech, prepare, speech to text, voices, search, offline packs, voice agents, live streaming and webhook checks, and return errors with a plain-language fix.

# JavaScript                                              # Python
npm install https://speakvora.com/sdk/speakvora-0.2.0.tgz   pip install https://speakvora.com/sdk/speakvora-0.2.0.tar.gz

const { Speakvora } = require("speakvora");               from speakvora import Speakvora
const sv = new Speakvora(process.env.SPEAKVORA_KEY);      sv = Speakvora(os.environ["SPEAKVORA_KEY"])
const clip = await sv.speak("Your order has shipped.");   clip = sv.speak("Your order has shipped.")
await sv.prepare(["Your table is ready."]);               sv.prepare(["Your table is ready."])
const { text } = await sv.listen(wavBuffer);              text = sv.listen(wav_bytes)["text"]
const turn = await sv.agent(id, agentKey).turn({ session: "c1", text: "Hours?" })

Single-file copies for pasting into a project: speakvora.js and speakvora.py. Keep your key on a server; for browsers mint a short-lived token with agent(id).sessions(). The npm and PyPI listings will follow the production launch.

Webhooks

Add a URL in the portal and we POST signed JSON events: test.ping, key.created, key.revoked. Each request carries x-speakvora-timestamp and x-speakvora-signature, the hex HMAC-SHA256 of timestamp + "." + raw body using your signing secret. Use Speakvora.verifyWebhook or verify_webhook from the SDKs, and reject anything older than five minutes.

Voices and languages

Use GET /api/v1/voices and GET /api/v1/languages for the live list. Voices that can speak any text live:

LanguagelangVoices
Englishenaf_heart, af_bella, af_nicole, af_sarah, af_sky, am_adam, am_michael
Spanishesef_dora, em_alex
Frenchfrff_siwis
Portugueseptpf_dora, pm_alex
Hindihihf_alpha, hm_omega
Japanesejajf_alpha, jm_kumo
Mandarin Chinesezhzf_001, zm_009

Other voices and languages play phrases from the library only, until live generation reaches them. The language you send must match the voice.

Limits

  • Text: 1 to 500 characters per request.
  • Rate: 15 requests per second per account on the free plan (100 on paid), counted over 10 seconds. Above that you get 429 rate_limited with retry_after_seconds.
  • Keys: up to 5 active keys per account. Webhooks: up to 3.
  • Numbers: English accepts digits. Every other language needs numbers written out in words.
  • Library first: library phrases are unlimited and instant. Brand-new catalog sentences are generated live and limited to 20 a minute per free account (120 paid), error live_rate_limited. Speed: library phrases and repeats are instant. A brand-new catalog sentence takes about 3 to 5 seconds, and the first one after a quiet period can take up to 30 seconds. A new custom-voice sentence takes about 2 seconds (styled: 4 to 10 seconds); a repeat is a cache hit (about 0.3 s).

Errors and fixes

Every error reply has error, message, fix and docs fields. The status codes are the standard ones.

StatuserrorMeaning and fix
401unauthorizedKey missing, wrong or revoked. Send it in x-api-key. Check for spaces or quotes around it.
400invalid_jsonBody is not valid JSON. Send content-type: application/json and check commas and quotes.
400text_required_max_500_charsText empty or over 500 characters. Split it.
400voice_requiredAdd a voice id.
422custom_voice_errorThe custom voice could not be used. Read detail: styles are English only, the voice must be ready, the id must be yours.
422style_not_availableStyles are not switched on for this environment.
400invalid_requestSpeech-to-text or streaming request malformed. Read detail.
404voice_not_liveThat voice cannot speak new text. Pick one from the table above.
422language_not_supported_for_voicelang does not match the voice. The reply says which language the voice speaks.
422language_not_supportedLanguage not available yet. The reply lists the ones that are.
422invalid_textText rejected. Read reasons_explained: usually digits in a non-English language, a placeholder like (name), or a safety claim.
404voice_not_available_for_textThe phrase exists, but not in that voice. The reply lists the voices that have it.
404not_in_libraryNot a library phrase and live generation could not run for it.
422generation_failed_qualityThe audio was too short or broken, so it was not served. Add a full stop or reword.
429(Too Many Requests)Over the rate limit. Wait a moment and retry with a short backoff.
502/503/504(gateway timeout)A new sentence took longer than the time limit, usually the first request after a quiet period. Retry the same request.

Troubleshooting checklist

  1. Run the quick check in the portal under Troubleshoot. It tests that the API is reachable and that your key works.
  2. Try the exact curl command from the portal in a terminal. If curl works but your app does not, the problem is in your app: the header name, an environment variable that is empty, or a proxy.
  3. Check the header. It is x-api-key, with no Bearer unless you use Authorization.
  4. Match language and voice. A Chinese voice with lang: "en" fails.
  5. Calling from a browser? Do not ship your key in public web code. Call Speakvora from your server, then give the audio link to the browser.
  6. Slow on the first call? New sentences are generated live. Repeats are instant. Retry once before assuming something is broken.
  7. Audio will not play? The reply is a link. Open the url directly. Browsers need a click or tap before they will autoplay audio.
  8. Still stuck? Paste the request and response in the portal's Troubleshoot tab. It checks your account and request and explains the cause.

FAQ

Why was a key rejected right after I created it? Check for a trailing space or newline when you pasted it. Keys start with sv_.

Can I send personal data? No. Use placeholders on your side and speak the personal part another way.

Are non-English voices final? They are labelled native_review_pending until a native speaker approves the language.

Privacy and retention

Two different things happen to audio, and the rules differ:

Catalog voices (/v1/speak, phrase library)Custom voices and agents
Who can fetch the audioAnyone with the link. Links are long and unguessable but public, and the same text in the same voice returns the same file.Only your account. Links are signed and last 5 minutes.
How long it is keptLibrary clips: served with a one-year cache lifetime. Live-generated clips: kept for reuse.75 days after the last time it was used; each reuse restarts the 75 days. Expired audio is refused and deleted.
Personal data in the textDo not send it. Send generic sentences and speak the personal part another way.Allowed if you have the right to use it. It stays private to your account.

Other things to know: conversation memory for agents is kept for 24 hours. Your text, and recordings you submit to make a voice, are processed by our AI provider (DeepInfra), which lists zero retention on the models we use; we do not keep your voice sample after the voice is made. You can delete one voice, one recording, or all private audio from the portal. Private audio has a storage quota (500 MB by default). Text with unresolved placeholders such as (name) or {name} is rejected, so build the real sentence on your side. Whether you may record, clone or process someone's voice or personal data is your responsibility under the laws that apply to you; we are not a law firm and this is not legal advice.

Safety note: Speakvora rejects text that is not a real sentence and never generates affirmative safety claims. Do not use synthetic speech to imply a person is on the line.

For AI assistants

Point your coding assistant at /docs/llms.txt and the OpenAPI spec; they contain everything needed to wire up the API.