How live voice generation works
What happens between your request and the audio link: library first, then generation, the checks that can reject text, and what to plan for in speed.
The path of a request
- Library lookup. If the exact sentence exists as a checked recording for the voice you asked for, the API returns that file's link immediately. No generation happens.
- Live generation. If it does not, and the voice supports live generation, the text is synthesised on a Kokoro voice, stored and returned as an MP3 link. The reply carries
live: true, thelanethat produced it and, for non-English text,review_status. - Reuse. The same sentence later is found in the store and returned quickly.
What you can ask for
12 live catalog voices across seven languages (English, Spanish, French, Portuguese, Hindi, Japanese, Mandarin). The language you send must match the voice. Text is 1 to 500 characters. English accepts digits; other languages need numbers written as words. Premium library voices play library sentences only.
Checks that can reject your text
| Code | Why | What to do |
|---|---|---|
| generation_failed_quality | The audio came out too short or broken, so it was not served | Add a full stop or reword |
| language_not_supported / cannot_detect_language | The language is not offered, or auto could not decide | Send an explicit lang |
| voice_not_available_for_text | That voice has no recording for this sentence and does not generate live | Choose a live voice |
| live_rate_limited | Too many brand-new sentences per minute on your account | Prepare ahead with /v1/prepare |
Speed, honestly
- Library lines and prepared sentences came back in about a tenth to a third of a second in our test.
- Brand-new sentences depend on whether a GPU worker is warm: about a second with one, several seconds without. The warm worker runs for paid accounts during certain hours; the speed page has the figures, the method and the caveats.
- Run
POST /v1/preparewith up to 200 sentences to have them ready before a customer needs them; preparing is not billed.
Quality and review
Non-English results are marked native_review_pending until a native speaker approves the language. English is our reference. We do not claim a quality score against other providers; listen on the voices page.
Streaming
The WebSocket speaks long text sentence by sentence and sends one audio link per sentence, so playback can start before the whole text is done. It returns links, not a raw audio stream. See the docs.
Next: text to speech · library plus live speech · your own voice · billing