Text to speech by API

Send up to 500 characters, get back a link to the audio. Lines already in the checked library come back instantly; any other text in a live voice is generated on demand. Billing is per character with no subscription.

Two ways to get speech

Library (catalog) speech

A checked sentence that already exists as audio. The API returns a link to the stored file, so there is no generation delay. Search the library with /v1/search.

Live generation

Any other text. It is synthesised on first request, then reused. Brand-new sentences took about 1 second in our latest test with a warm GPU worker, and several seconds without one (see the speed table). Plan for both.

Voices and languages today

The API lists 12 live catalog voices across 7 languages (English, Spanish, French, Hindi, Japanese, Portuguese, Chinese). Live generation is currently offered for Kokoro voices; the language you send must match the voice. Numbers: English accepts digits, other languages need numbers written as words. Non-English results carry a "native review pending" label until a speaker approves them.

VoiceIdTierDescriptionChecked clipsStatus
Heartaf_heartKokoroWarm, clear and conversational. Our reference voice.en: 7,730live
Easygoing Friendeasygoing-friendPremiumRelaxed, natural and friendly. One identity across many languages.en: 802, zh: 15library only
Doraef_doraKokoroSpanishes: 29live
Alexem_alexKokoroSpanishes: 31live
Siwisff_siwisKokoroFrenchfr: 5live
Dorapf_doraKokoroBrazilian Portuguesept: 29live
Alexpm_alexKokoroBrazilian Portuguesept: 28live
Alphahf_alphaKokoroHindihi: 16live
Omegahm_omegaKokoroHindihi: 18live
Alphajf_alphaKokoroJapaneseja: 19live
Kumojm_kumoKokoroJapaneseja: 20live
Mandarin F1zf_001KokoroMandarin Chinesezh: 32live
Mandarin M1zm_009KokoroMandarin Chinesezh: 32live
LanguageCodeLibrary clipsReview
Englishen7,730approved
Spanishes31pending native-speaker review
Frenchfr5pending native-speaker review
Hindihi19pending native-speaker review
Japaneseja23pending native-speaker review
Portuguesept30pending native-speaker review
Chinesezh33pending native-speaker review

Price

WhatList priceUnit
Catalog voices (Kokoro)$0.70per 1,000,000 characters
Your own voice$6per 1,000,000 characters
Styled voice (private preview, off in production)$17per 1,000,000 characters
Free allowance100,000 charactersper account, no card

Library lines are billed by character like any other text, and a repeat of the same sentence is billed again. Early-access list prices; List prices. See pricing and .

Speed

Library and prepared phrases come back in about a tenth to a third of a second in our test. Brand-new sentences depend on whether a GPU worker is warm, which is limited to certain hours for paid accounts. See the measured table, the method and the caveats on the speed page. These are our measurements, not guarantees.

First request

cURL

curl -X POST https://speakvora.com/api/v1/speak \
  -H "x-api-key: $SPEAKVORA_KEY" \
  -H "content-type: application/json" \
  -d '{"text":"Your order is ready for pickup.","voice":"af_heart","lang":"en"}'

# 200 {"url":"https://speakvora.com/voice/v2/af_heart/..../r1.m4a","lang":"en","voice":"af_heart","characters":32,"format":"m4a"}

JavaScript (server)

// Run this on your server (Node 18+). Never put the key in browser code.
const r = await fetch("https://speakvora.com/api/v1/speak", {
  method: "POST",
  headers: { "x-api-key": process.env.SPEAKVORA_KEY, "content-type": "application/json" },
  body: JSON.stringify({ text: "Your order is ready for pickup.", voice: "af_heart", lang: "en" }),
});
if (!r.ok) throw new Error((await r.json()).error);   // e.g. unauthorized, free_allowance_used, rate_limited
const { url } = await r.json();                        // link to the audio file

Python (requests)

import os, requests

r = requests.post(
    "https://speakvora.com/api/v1/speak",
    headers={"x-api-key": os.environ["SPEAKVORA_KEY"]},
    json={"text": "Your order is ready for pickup.", "voice": "af_heart", "lang": "en"},
    timeout=30,
)
r.raise_for_status()
print(r.json()["url"])   # link to the audio file

Responses are JSON with a url. Fetch that URL to download the audio. Limits: 1–500 characters per request; 15 requests per second on the free plan (100 paid), with 429 rate_limited and retry_after_seconds when exceeded. Retry 502/503/504 once for brand-new sentences.

Browser use

Plain /v1/speak keys are not scoped for browsers. Call the API from your server and send the browser the returned audio link. Browser-safe session tokens exist for voice agents only (agents).

Next

Add speech to a web app · Industry packs · Your own voice · Streaming speech