Speech to text, $0.0015 a minute
Send audio as a file or stream it live over a WebSocket and get text back. Your first 60 minutes are free.
What it does
- File transcription:
POST /v1/listenwith base64 audio (wav, mp3, m4a, ogg, flac, webm; 1 KB to 4.5 MB per request). - Live transcription: WebSocket with automatic end-of-speech detection; the final text arrives about 0.8 seconds after speech stops in our test (one observation, not a guarantee).
- Billing: per second of audio, listed at $0.0015 per minute.
What it does not do yet
No speaker labels and no word timestamps. Long recordings must be sent in pieces. We have not published an accuracy benchmark; test it on your own audio.
Example
cURL
curl -X POST https://speakvora.com/api/v1/listen -H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
-d "{\"audio_base64\":\"$(base64 < clip.wav | tr -d '\n')\",\"format\":\"wav\",\"language\":\"en\"}"
# 200 {"text":"...","seconds":3,"language":"en"}Speed
Measured transcription times for 4 and 15 second clips are on the speed page, with the method. They are our measurements, not guarantees.
Price in native units
| Provider | Model | Price | Unit |
|---|---|---|---|
| Speakvora | Speech to text | $0.0015 | per minute (billed per second) |
| OpenAI | whisper | $0.006 | per minute |
| Deepgram | Nova-3 monolingual, pre-recorded | $0.0043 (promotional; regular $0.0077 listed) | per minute |
| ElevenLabs | Scribe v2 | $0.22 | per hour (about $0.0037 per minute) |
Sources and checked date on the (checked 2026-10-08). Models differ in accuracy, languages and features, so this is a price comparison only.