Stream live transcription over a WebSocket
Authenticate, stream 16 kHz PCM audio in short frames, and read final text shortly after speech stops.
The WebSocket takes JSON text frames. The first frame must authenticate. The server detects the end of speech by loudness and replies with final text.
Message flow
address: wss://48f8pzwtee.execute-api.us-east-1.amazonaws.com/v1
> {"type":"auth","key":"YOUR_KEY"} < {"type":"ready"}
> {"type":"listen.start","language":"en","rate":16000,"interim":false} # optional: endpoint_ms (700), vad_threshold (400)
> {"type":"audio","seq":0,"data":"<base64 PCM16 mono, 100-500 ms>"} # seq counts up from 0
< {"type":"final","text":"...","seconds":4,"latency_ms":830}
> {"type":"listen.end"} # flushNode 22+ (global WebSocket), server side only
const ws = new WebSocket("wss://48f8pzwtee.execute-api.us-east-1.amazonaws.com/v1");
let seq = 0;
ws.onopen = () => ws.send(JSON.stringify({ type: "auth", key: process.env.SPEAKVORA_KEY }));
ws.onmessage = (m) => {
const msg = JSON.parse(m.data);
if (msg.type === "ready") {
ws.send(JSON.stringify({ type: "listen.start", language: "en", rate: 16000, interim: false }));
// send each 100-500 ms chunk of PCM16 mono as base64:
// ws.send(JSON.stringify({ type: "audio", seq: seq++, data: chunkBase64 }));
}
if (msg.type === "final") console.log(msg.text, msg.latency_ms + " ms");
};Tuning
- End-of-speech detection is loudness-based: raise
vad_thresholdin noisy rooms. - With
interim: trueyou also get partial results about every two seconds of speech; they re-transcribe the buffer and are billed as audio seconds. - The final text arrived about 0.8 seconds after speech stopped in our earlier test; the speed page has the method. It is a measurement, not a guarantee.
- Keep the key on your server and stream from there; never put it in a public web page.