Turn text into speech with one curl request

Make a single POST request to /v1/speak with text and a voice ID to get a URL to audio. Authenticate with x-api-key header.

Short answer. Send a POST request to /v1/speak with your text, a voice ID, and your API key in the x-api-key header. The API returns a URL to the audio file in M4A or MP3 format. Fetch the URL to play or download the audio.

The basic request

A single curl command turns text into speech. You send JSON with the text, a voice ID, and a language code. The API returns a URL to audio that you can play or download.

curl example

curl -X POST https://speakvora.com/api/v1/speak \
  -H 'x-api-key: your-key-here' \
  -H 'Content-Type: application/json' \
  -d '{"text": "Hello world", "voice": "af_heart", "lang": "en"}'

The response

The API replies with a JSON object containing the audio URL and metadata.

Response example

{
  "url": "https://speakvora.com/...",
  "lang": "en",
  "voice": "af_heart",
  "characters": 11,
  "format": "m4a",
  "live": true
}

Request fields

  • text: 1 to 500 characters. Required.
  • voice: voice ID like af_heart or em_alex. Required.
  • lang: language code like en, es, fr, pt, hi, ja, or zh. Optional; defaults to en.

Response fields

  • url: public, unguessable URL to the audio file. Fetch it to play or save.
  • format: m4a or mp3.
  • characters: number of characters billed (the length of your text).
  • live: true if the text was generated on demand; false if it was already in the library.
  • lang and voice: echo of your request.

Playing the audio

Once you have the URL, fetch it with GET and play or save the file. Library phrases are instant; new text is generated live (about 3 to 5 seconds, up to 30 seconds after idle).

Fetch and play

curl -o speech.m4a 'https://speakvora.com/...'
# Then play speech.m4a with your audio player

Authentication

Include your API key in the x-api-key header. Keep your key on a server; do not expose it in client-side code. You can also use the Authorization header with Bearer token format.

Billing and limits

  • Billed per character of text. Catalog voices cost $0.70 per 1M characters.
  • Every account gets 100,000 free characters.
  • Text must be 1 to 500 characters per request.
  • English accepts digits; other languages need numbers written as words.
  • Do not send personal data through catalog voices; audio is stored at public URLs.

Frequently asked questions

How long does it take to get audio?

Library phrases (text already in the catalog) return instantly. New text is generated live, typically 3 to 5 seconds, up to 30 seconds after idle.

What voices can I use?

Catalog voices like af_heart (Heart, female, English), em_alex (Alex, male, Spanish), ff_siwis (Siwis, female, French), and others. Call GET /v1/voices to list all available voices and their languages.

Can I use the same text twice without being billed twice?

Yes. If the text is already in the library, the second request returns the cached audio instantly and is billed the same as the first. The response field live tells you whether the audio was generated or cached.

What if my text contains a placeholder like (name)?

Unresolved placeholders are rejected with error invalid_text. Build the real sentence yourself before sending it.

Can I stream audio back to the browser?

The /v1/speak endpoint returns a URL to audio. For true streaming, use the WebSocket API with the speak message type.

Related: API documentation, pricing, developer guides and more answers.