How live voice generation works

What happens between your request and the audio link: library first, then generation, the checks that can reject text, and what to plan for in speed.

The path of a request

  1. Library lookup. If the exact sentence exists as a checked recording for the voice you asked for, the API returns that file's link immediately. No generation happens.
  2. Live generation. If it does not, and the voice supports live generation, the text is synthesised on a Kokoro voice, stored and returned as an MP3 link. The reply carries live: true, the lane that produced it and, for non-English text, review_status.
  3. Reuse. The same sentence later is found in the store and returned quickly.

What you can ask for

12 live catalog voices across seven languages (English, Spanish, French, Portuguese, Hindi, Japanese, Mandarin). The language you send must match the voice. Text is 1 to 500 characters. English accepts digits; other languages need numbers written as words. Premium library voices play library sentences only.

Checks that can reject your text

CodeWhyWhat to do
generation_failed_qualityThe audio came out too short or broken, so it was not servedAdd a full stop or reword
language_not_supported / cannot_detect_languageThe language is not offered, or auto could not decideSend an explicit lang
voice_not_available_for_textThat voice has no recording for this sentence and does not generate liveChoose a live voice
live_rate_limitedToo many brand-new sentences per minute on your accountPrepare ahead with /v1/prepare

Speed, honestly

  • Library lines and prepared sentences came back in about a tenth to a third of a second in our test.
  • Brand-new sentences depend on whether a GPU worker is warm: about a second with one, several seconds without. The warm worker runs for paid accounts during certain hours; the speed page has the figures, the method and the caveats.
  • Run POST /v1/prepare with up to 200 sentences to have them ready before a customer needs them; preparing is not billed.

Quality and review

Non-English results are marked native_review_pending until a native speaker approves the language. English is our reference. We do not claim a quality score against other providers; listen on the voices page.

Streaming

The WebSocket speaks long text sentence by sentence and sends one audio link per sentence, so playback can start before the whole text is done. It returns links, not a raw audio stream. See the docs.

Next: text to speech · library plus live speech · your own voice · billing