Transcribe an audio file with the API
Send a recording as base64, get text and the billed seconds back, and handle the size limit and errors.
The file endpoint takes one request per recording up to 4.5 MB. Longer audio goes in pieces.
Request
cURL
B64=$(base64 < clip.wav | tr -d '\n')
curl -X POST https://speakvora.com/api/v1/listen \
-H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
-d "{\"audio_base64\":\"$B64\",\"format\":\"wav\",\"language\":\"en\"}"
# 200 {"text":"Your refund will arrive in three to five business days.","seconds":3,"language":"en"}Python (requests)
import base64, os, requests
audio = base64.b64encode(open("clip.wav", "rb").read()).decode()
r = requests.post("https://speakvora.com/api/v1/listen", headers={"x-api-key": os.environ["SPEAKVORA_KEY"]},
json={"audio_base64": audio, "format": "wav", "language": "en"}, timeout=60)
r.raise_for_status()
print(r.json()["text"])Details that matter
- Formats: wav, mp3, m4a, ogg, flac, webm. Size: 1 KB to 4.5 MB per request.
- Language is optional.
- Billing: per second of audio, listed at $0.0015 per minute; the first 60 minutes per account are free.
- Errors:
400 invalid_request(format, size or base64),402 free_allowance_used,502 upstream_unavailable(retry once). - Not offered yet: speaker labels and word timestamps.
For long recordings, split at silence, send each piece and join the text. Test accuracy on your own audio first; we publish no accuracy benchmark. See the price comparison and live transcription.