WebSocket Streaming: Audio Links and Live Transcription
Speakvora streams speech and transcription over WebSocket with audio returned as URLs per sentence, not raw bytes. Live transcription and agent turns supported.
What streaming does Speakvora offer?
Speakvora provides WebSocket streaming for two use cases: live transcription (speech to text) and streamed speech generation. Audio is not returned as raw bytes; instead, each sentence or transcription result arrives as a URL you can play. This design keeps latency low and lets you start playing audio before the entire response is complete.
Connecting and authenticating
Connect to the WebSocket endpoint and send an authentication message as your first frame. The server replies when ready.
WebSocket authentication
First message (JSON text frame):
{"type": "auth", "key": "your-api-key"}
Server reply:
{"type": "ready"}Live transcription (listen.start flow)
Start a transcription session by sending listen.start with your language and audio settings. Then send audio frames as base64-encoded PCM16 mono data. The server detects speech endpoints (silence) and sends final results when speech stops. Interim results are available if you request them.
- Audio frames: 100–500 ms each, PCM16 mono, base64-encoded
- Endpoint detection: speech is considered complete after endpoint_ms of silence
- VAD threshold: raise vad_threshold in noisy rooms (default 400)
- Interim results: set interim to true to receive partial transcriptions; these are re-transcriptions of the buffer and are billed as audio seconds
- Final result: arrives about 0.8 s after speech ends
Transcription example
{"type": "listen.start", "language": "en", "rate": 16000, "interim": false, "endpoint_ms": 700, "vad_threshold": 400}
{"type": "audio", "seq": 0, "data": "<base64 PCM16 mono>"}
{"type": "audio", "seq": 1, "data": "<base64 PCM16 mono>"}
{"type": "listen.end"}
Server sends:
{"type": "final", "text": "...", "seconds": 2.5, "latency_ms": 800}Streamed speech generation
Request speech for one or more sentences. The server sends audio URLs as each sentence is ready, in order. Custom voices (cv_...) work alongside catalog voices.
Streamed speech request
{"type": "speak", "id": "1", "text": "Hello world", "voice": "af_heart", "lang": "en"}
Server sends (one per sentence, in order):
{"type": "audio", "index": 0, "text": "Hello world", "url": "...", "source": "library|generated|cache"}
{"type": "done"}Streamed voice agents
Send agent.turn to run a turn and receive the agent's response as it is generated. Transcript, text, and audio URLs arrive as separate events. Use an agent session token (st.…) to authenticate for that agent only.
- Agent turns are turn-based: the response arrives when the turn is complete, typically 2–4 s with a custom voice
- Session tokens are valid for 10 minutes
- Audio URLs arrive in order as each sentence is ready
What is not streamed
Audio is always returned as URLs to play, not raw bytes. Word timestamps and speaker labels are not available. True incremental decoding and raw audio streaming back are not offered. For longer recordings, send audio in pieces and handle each transcription result separately.
Billing
Interim transcription results are billed as audio seconds. Final results are billed once per request. Streamed speech is billed per character as usual. Agent turns are billed per turn.
Frequently asked questions
Does Speakvora return raw audio bytes over WebSocket?
No. Audio is returned as URLs to play, not raw streamed bytes. This keeps latency low and lets you start playing before the entire response is ready.
How do I start a live transcription?
Send {"type": "listen.start", "language": "en", "rate": 16000, "interim": false, "endpoint_ms": 700, "vad_threshold": 400}, then send audio frames with seq counting up from 0, then send {"type": "listen.end"}. The server sends final results when speech stops.
Are interim transcription results billed?
Yes. Interim results are re-transcriptions of the buffer and are billed as audio seconds. Final results are billed once per request.
Can I use an agent session token to authenticate on WebSocket?
Yes. An agent session token (st.…) authenticates for that agent only and is valid for 10 minutes.
Related: API documentation, pricing, developer guides and more answers.