Text to speech API speed and price, with sources
List price per million characters and time to first audio for four providers. Vendor-stated model latency is kept apart from independent measurements, because they do not measure the same thing.
| Provider | List price per 1M characters | Stated first audio | Independent or our own measurement |
|---|---|---|---|
| ElevenLabs (Flash v2.5) | $40 to $80 | about 75 ms, excludes network (ElevenLabs models page) | 197 ms median, network included (Vapi Humanness Index) |
| Cartesia (Sonic) | about $38 to $50, credit based | under 90 ms, model latency (Cartesia) | 128 to 159 ms median, network included (Vapi Humanness Index) |
| Deepgram (Aura-2) | $15 to $30 | about 90 ms steady state, under 200 ms at launch (Deepgram changelog) | not found |
| Speakvora | $3 (your own voice $6) | 82 ms short sentence on our larger GPU, 126 ms on our standard GPU, measured inside our cloud region (our speed page) | 0.54 s to the first audio sentence over our WebSocket, measured from a home connection in New York |
How to read this
- Vendor-stated latency is usually model time only and leaves out the network. Independent figures include the network from the tester.
- Our 82 to 126 ms is measured from inside the same cloud region as our API, so it has almost no network in it. The 0.54 s includes a home connection, which is why the two differ.
- Prices are official list prices at the standard tier, re-checked monthly, with each source on the comparison page.
What we do not claim
We do not claim to be faster than these providers. We have not benchmarked voice quality against them, and we measured only our own service. Their figures are quoted from the sources linked above and may have changed since.
Where Speakvora is ahead
- Price: $3 per million characters for catalog voices.
- Offline bundles: one-time packs that play with no per-play fee and no generation delay.
- Stored sentences: a sentence we have already generated comes back in about 0.14 s.
Questions
Is Speakvora faster than ElevenLabs, Cartesia or Deepgram?
No, not on first audio. Their published time to first audio is lower than the 0.54 s we measure over our streaming WebSocket from a home connection. Inside our own cloud region our server returns first audio in 82 to 126 ms for a short sentence, which is in the same range as their stated model latency. We compete on price and on offline bundles, not on being the fastest.
How much cheaper is Speakvora?
Catalog voices are $3 per million characters. The providers above list $15 to $80 per million at the standard tier we compared. Official prices and their sources are on the comparison page.
What are offline bundles?
One-time packs of checked prompts (menus, confirmations, errors) delivered as audio files. They play from the device with no per-play fee, which is the fastest option because nothing is generated.