# Speakvora Voice API (for AI coding assistants) Speakvora turns text into speech. Base URL (development): https://speakvora.com/api . Auth: header `x-api-key: ` (or `Authorization: Bearer `). Keep keys on a server. Every account gets 100,000 free characters (402 free_allowance_used after that until billing opens). Errors have: error, message, fix, docs. ## Speak with a catalog voice POST /v1/speak {"text": "...", "voice": "af_heart", "lang": "en"} -> 200 {"url": "...", "lang", "voice", "characters", "format": "m4a|mp3", "live": true?} Then GET the url (audio). Library phrases are instant; other text is generated live (about 3-5 s, up to 30 s after idle). Text: 1-500 chars. Live catalog voices: en af_heart af_bella af_nicole af_sarah af_sky am_adam am_michael; es ef_dora em_alex; fr ff_siwis; pt pf_dora pm_alex; hi hf_alpha hm_omega; ja jf_alpha jm_kumo; zh zf_001 zm_009. lang must match the voice. English accepts digits; every other language needs numbers written as words. Unresolved placeholders like (name) are rejected (error invalid_text): build the real sentence yourself. Catalog audio is stored at public, unguessable URLs and may be reused for the same text. DO NOT send personal data through catalog voices. ## Other endpoints GET /v1/voices, GET /v1/languages, GET /v1/packs (offline packs), GET /v1/search?q=...&lang=en&limit=10 (find library phrases by words), GET /api/status. No-signup try-it: POST /demo/speak {"text" (<=140 chars), "voice"} for 9 voices, limited per visitor. MCP server (streamable HTTP, JSON-RPC): POST /mcp with x-api-key. Tools: speak, list_voices, search_phrases. Claude Code: claude mcp add --transport http speakvora https://speakvora.com/api/mcp --header "x-api-key: KEY" ## Custom voices (portal) and voice agents Customers make voices in the portal (https://speakvora.com/dashboard): record a paragraph, upload 10-60 s of speech (needs confirmation it is their voice or they have permission), or describe and tune a voice with meters (gender, age, pitch, pace, warmth, energy). Max 3 per account. A voice is "ready" only after a verification (we speak a test sentence and listen). Custom voices: English supported; French and German beta; Spanish, Hindi, Chinese, Japanese NOT offered (not intelligible when tested). Laughter tags exist only for English custom voices and are off by default. Delivery styles (happy, excited, sorry, calm, serious) exist for English; pitch and speed are not controllable. Agents are built in the portal's Agent builder (chat + editable config). "Generate API" publishes a numbered version and returns the endpoint and an agent-scoped key (shown once; works only for that agent). POST /v1/agents/{agent_id}/turn {"session": "call-1", "text": "..."} or {"audio_b64": ""} or {"start": true} -> {"reply", "generation", "sentences": [{"index","text","audio_url","source": "library|cache|generated|live"}], "timings": {...}, "session", "agent", "version"} Play sentences[].audio_url in order. Browsers: server calls POST /v1/agents/{id}/sessions (agent key) -> {"token"} valid 10 minutes; browser sends Authorization: Bearer . Tokens/agent keys reach only that agent. Interruption: send "interrupted": {"generation": "...", "played": N} on the next turn, or POST /v1/agents/{id}/interrupt {session, generation, played}. Memory then keeps only what was played. Limits: turn-based (response arrives when the turn is complete, typically 2-4 s with a custom voice); for live streaming use the WebSocket API below (agent.turn); no phone numbers. 200 free turns, then $10 per 1,000 turns (early access). ## Own voice speech and styles (POST /v1/speak with a custom voice) POST /v1/speak {"text","voice":"cv_...","style":"neutral|happy|excited|sorry|calm|serious"(optional, English only)} -> {"url" (signed, 5 minutes), "format": "mp3|wav", "source": "generated|cache", "style", "style_applied", "engine"} Styled speech is WAV, $17 per 1M characters; plain own-voice speech is $6. If style_applied is false the backup engine answered without the style and the plain rate applies. A custom-voice sentence is stored only after its second request; stored audio is private and deleted after 75 days without use. Accents cannot be requested. Prices per 1M characters: catalog $0.70, own voice $6, styled $17. Speech to text $0.0015/min. Agents $10 per 1,000 turns. Free: 100,000 characters, 60 STT minutes, 200 agent turns. ## Offline packs GET /v1/packs -> {"packs":[{"id","title","lang","voice","phrases","bytes","price_usd","sample_url","examples"}]}. Free 25-phrase sample at /packs/{id}/sample.zip. Full pack is a one-time card purchase on /packs; then GET /v1/packs/{id}/download (x-api-key) -> {"url"} (zip, valid 5 minutes), or 402 pack_not_purchased. Zip: manifest.json, phrases.csv, audio/NNNNN.m4a, LICENSE.txt. ## Speech to text POST /v1/listen {"audio_base64": "", "format": "wav|mp3|m4a|ogg|flac|webm", "language": "en" (optional)} -> 200 {"text", "seconds", "language"} Audio 1 KB to 4.5 MB per request (send longer recordings in pieces). Billed per second of audio, rounded up per request. 60 free minutes per account. Errors: 400 invalid_request, 402 free_allowance_used, 502 upstream_unavailable. ## Live streaming (WebSocket) Connect to wss://48f8pzwtee.execute-api.us-east-1.amazonaws.com/v1 (dev; a custom domain comes later). JSON text frames; the first message must be {"type":"auth","key":"..."} -> {"type":"ready"}. Live transcription: {"type":"listen.start","language":"en","rate":16000,"interim":false,"endpoint_ms":700,"vad_threshold":400}, then {"type":"audio","seq":0,"data":""} with seq counting up from 0 (100-500 ms per frame), then {"type":"listen.end"} to flush. Server sends {"type":"final","text","seconds","latency_ms"} when speech stops for endpoint_ms (or 25 s), and {"type":"partial",...} about every 2 s of speech when interim is true (interim results are re-transcriptions of the buffer and are billed as audio seconds). Endpointing is an energy detector: raise vad_threshold in noisy rooms. Utterance results arrive about 0.8 s after the end of speech (measured). Streamed speech: {"type":"speak","id":"1","text":"...","voice":"af_heart","lang":"en"} -> one {"type":"audio","index","text","url","source"} per sentence, in order, as each is ready, then {"type":"done"}. Custom voice ids (cv_...) work too. Streamed agent: {"type":"agent.turn","agent":"","text":"...","session":"s1"} -> the same events as the turn engine (transcript, text, audio, done), pushed as they happen. An agent session token (st.…) can authenticate for this agent only. Not offered: word timestamps, speaker labels, true incremental decoding, raw audio streaming back (audio arrives as URLs to play). ## Privacy and retention Custom-voice and agent audio is private to the account, linked with signed 5-minute URLs, kept 75 days after last use (each reuse restarts the 75 days), then deleted. Agent conversation memory: 24 hours. Voice samples are not kept after the voice is made. Text is processed by DeepInfra. Delete one voice, one recording or all private audio in the portal. Storage quota 500 MB default. Public library clips are served with a one-year cache lifetime. Do not present the voice as a human; disclose automated voices where law or good practice requires. Safety-critical alerts must not depend on the network. ## SDKs and webhooks Single-file clients: https://speakvora.com/sdk/speakvora.js and /sdk/speakvora.py. Webhooks (portal): events test.ping, key.created, key.revoked; header x-speakvora-signature = hex HMAC-SHA256(secret, timestamp + "." + rawBody), header x-speakvora-timestamp; reject if older than 5 minutes. Troubleshooting: portal > Troubleshoot, and https://speakvora.com/docs#errors .