OpenAI text to speech alternative: lower per-character and per-minute cost

OpenAI's speech endpoint is convenient because it sits beside the language models. If voice cost has become a line item, a cheaper dedicated speech service is worth testing against it.

Who switches

  • Voice is a growing share of your API bill
  • You transcribe a lot of audio with Whisper
  • You want a finished phrase library for kiosks and IVR

What you gain

  • $0.70 per million characters against $15 (tts-1) and $30 (tts-1-hd)
  • $0.0015 per minute transcription against $0.006 (Whisper)
  • A checked English phrase library, industry packs (47 live across 23 industries today, more being built for every industry)
  • Your own voice at $6 per million

What you give up

  • Instruction-steered delivery from gpt-4o-mini-tts
  • Streaming audio responses
  • One account for language and speech
  • Larger voice and language coverage

The code change

Request shapes were taken from each provider's own documentation on 2026-10-08. Provider APIs change; check theirs before you migrate.

OpenAI text to speech (docs: developers.openai.com/api/docs/guides/text-to-speech)

curl https://api.openai.com/v1/audio/speech \
  -H "Authorization: Bearer $OPENAI_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini-tts","input":"Hello, how can I help you today?","voice":"marin","response_format":"mp3"}' \
  --output prompt.mp3
# response body is the audio bytes

Speakvora text to speech

curl -X POST https://speakvora.com/api/v1/speak \
  -H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
  -d '{"text":"Hello, how can I help you today?","voice":"af_heart","lang":"en"}'
# JSON reply: {"url":"https://speakvora.com/...m4a","lang":"en","voice":"af_heart","characters":32,"format":"m4a"}
curl -s "$URL" --output prompt.m4a        # second step: download (or play) the link

What changes in practice: OpenAI takes `model`, `input` and `voice` and returns audio bytes; its newest model also takes an `instructions` field to steer delivery. Speakvora takes `text`, `voice` and `lang` and returns a link, and has no equivalent of `instructions` in production. The text limit is 500 characters per request, and Speakvora returns JSON with a url that you fetch or play. Library links are stable public files; custom-voice links last about five minutes.

A safe migration order

  1. Keep OpenAI for language and delivery-sensitive speech.
  2. Move fixed prompts and notifications to Speakvora library lines or /v1/speak.
  3. Move bulk transcription after testing accuracy on your recordings.
  4. Keep a fallback to OpenAI behind a flag during the trial.

Is it actually cheaper?

OpenAI's character-billed models are tts-1 at $15 and tts-1-hd at $30 per million characters. Speakvora's catalog voices are $0.70, about 95% and 98% lower at list price. OpenAI's gpt-4o-mini-tts is billed in tokens ($0.60 per million text-input tokens and $12 per million audio-output tokens), and we do not convert tokens to characters because audio output tokens depend on the length of speech generated.

Full numbers, scenarios and sources: Speakvora vs OpenAI.

Test these before you commit

  • Listen to your ten most-spoken lines on both, and note whether you used `instructions` to shape them.
  • Re-measure your transcription accuracy on your own recordings before moving Whisper workloads.
  • Check that you do not depend on streamed audio chunks; Speakvora returns a link, and streams sentence links over its WebSocket.
  • Keep OpenAI as a fallback behind a flag for two weeks.

Speakvora is in early access with billing is live; questions to support@eyeguide.ai. Latency: see the speed page.