Text to speech by API
Send up to 500 characters, get back a link to the audio. Lines already in the checked library come back instantly; any other text in a live voice is generated on demand. Billing is per character with no subscription.
Two ways to get speech
Library (catalog) speech
A checked sentence that already exists as audio. The API returns a link to the stored file, so there is no generation delay. Search the library with /v1/search.
Live generation
Any other text. It is synthesised on first request, then reused. Brand-new sentences took about 1 second in our latest test with a warm GPU worker, and several seconds without one (see the speed table). Plan for both.
Voices and languages today
The API lists 12 live catalog voices across 7 languages (English, Spanish, French, Hindi, Japanese, Portuguese, Chinese). Live generation is currently offered for Kokoro voices; the language you send must match the voice. Numbers: English accepts digits, other languages need numbers written as words. Non-English results carry a "native review pending" label until a speaker approves them.
| Voice | Id | Tier | Description | Checked clips | Status |
|---|---|---|---|---|---|
| Heart | af_heart | Kokoro | Warm, clear and conversational. Our reference voice. | en: 7,730 | live |
| Easygoing Friend | easygoing-friend | Premium | Relaxed, natural and friendly. One identity across many languages. | en: 802, zh: 15 | library only |
| Dora | ef_dora | Kokoro | Spanish | es: 29 | live |
| Alex | em_alex | Kokoro | Spanish | es: 31 | live |
| Siwis | ff_siwis | Kokoro | French | fr: 5 | live |
| Dora | pf_dora | Kokoro | Brazilian Portuguese | pt: 29 | live |
| Alex | pm_alex | Kokoro | Brazilian Portuguese | pt: 28 | live |
| Alpha | hf_alpha | Kokoro | Hindi | hi: 16 | live |
| Omega | hm_omega | Kokoro | Hindi | hi: 18 | live |
| Alpha | jf_alpha | Kokoro | Japanese | ja: 19 | live |
| Kumo | jm_kumo | Kokoro | Japanese | ja: 20 | live |
| Mandarin F1 | zf_001 | Kokoro | Mandarin Chinese | zh: 32 | live |
| Mandarin M1 | zm_009 | Kokoro | Mandarin Chinese | zh: 32 | live |
| Language | Code | Library clips | Review |
|---|---|---|---|
| English | en | 7,730 | approved |
| Spanish | es | 31 | pending native-speaker review |
| French | fr | 5 | pending native-speaker review |
| Hindi | hi | 19 | pending native-speaker review |
| Japanese | ja | 23 | pending native-speaker review |
| Portuguese | pt | 30 | pending native-speaker review |
| Chinese | zh | 33 | pending native-speaker review |
Price
| What | List price | Unit |
|---|---|---|
| Catalog voices (Kokoro) | $0.70 | per 1,000,000 characters |
| Your own voice | $6 | per 1,000,000 characters |
| Styled voice (private preview, off in production) | $17 | per 1,000,000 characters |
| Free allowance | 100,000 characters | per account, no card |
Library lines are billed by character like any other text, and a repeat of the same sentence is billed again. Early-access list prices; List prices. See pricing and .
Speed
Library and prepared phrases come back in about a tenth to a third of a second in our test. Brand-new sentences depend on whether a GPU worker is warm, which is limited to certain hours for paid accounts. See the measured table, the method and the caveats on the speed page. These are our measurements, not guarantees.
First request
cURL
curl -X POST https://speakvora.com/api/v1/speak \
-H "x-api-key: $SPEAKVORA_KEY" \
-H "content-type: application/json" \
-d '{"text":"Your order is ready for pickup.","voice":"af_heart","lang":"en"}'
# 200 {"url":"https://speakvora.com/voice/v2/af_heart/..../r1.m4a","lang":"en","voice":"af_heart","characters":32,"format":"m4a"}JavaScript (server)
// Run this on your server (Node 18+). Never put the key in browser code.
const r = await fetch("https://speakvora.com/api/v1/speak", {
method: "POST",
headers: { "x-api-key": process.env.SPEAKVORA_KEY, "content-type": "application/json" },
body: JSON.stringify({ text: "Your order is ready for pickup.", voice: "af_heart", lang: "en" }),
});
if (!r.ok) throw new Error((await r.json()).error); // e.g. unauthorized, free_allowance_used, rate_limited
const { url } = await r.json(); // link to the audio filePython (requests)
import os, requests
r = requests.post(
"https://speakvora.com/api/v1/speak",
headers={"x-api-key": os.environ["SPEAKVORA_KEY"]},
json={"text": "Your order is ready for pickup.", "voice": "af_heart", "lang": "en"},
timeout=30,
)
r.raise_for_status()
print(r.json()["url"]) # link to the audio fileResponses are JSON with a url. Fetch that URL to download the audio. Limits: 1–500 characters per request; 15 requests per second on the free plan (100 paid), with 429 rate_limited and retry_after_seconds when exceeded. Retry 502/503/504 once for brand-new sentences.
Browser use
Plain /v1/speak keys are not scoped for browsers. Call the API from your server and send the browser the returned audio link. Browser-safe session tokens exist for voice agents only (agents).
Next
Add speech to a web app · Industry packs · Your own voice · Streaming speech