Amazon Polly alternative: API-key simplicity and a ready-made phrase library
Polly is a good service. People look for an alternative when they leave AWS, when they want fewer moving parts than IAM and request signing, or when they would rather buy finished prompts than synthesise them.
Who switches
- You are not on AWS and do not want to create an account and IAM policy just for speech
- You synthesise the same prompts repeatedly and would rather download them
- You exceed Polly's free tier and want a lower paid rate
What you gain
- An API key header instead of AWS credentials
- $0.70 per million against $4 Standard and $16 Neural
- A checked phrase library, industry packs (47 live across 23 industries today, more being built for every industry)
- Your own voice at $6 per million
What you give up
- Polly's free tier of 5 million Standard characters a month
- More languages, voices and SSML
- AWS integration and SLAs
- A stated right to cache output on its pricing page
The code change
Request shapes were taken from each provider's own documentation on 2026-10-08. Provider APIs change; check theirs before you migrate.
Amazon Polly (AWS SDK for Python; credentials come from your AWS configuration, not an API key header)
import boto3
polly = boto3.client("polly")
r = polly.synthesize_speech(Text="Hello, how can I help you today?", OutputFormat="mp3", VoiceId="Joanna", Engine="neural")
open("prompt.mp3", "wb").write(r["AudioStream"].read())Speakvora text to speech
curl -X POST https://speakvora.com/api/v1/speak \
-H "x-api-key: $SPEAKVORA_KEY" -H "content-type: application/json" \
-d '{"text":"Hello, how can I help you today?","voice":"af_heart","lang":"en"}'
# JSON reply: {"url":"https://speakvora.com/...m4a","lang":"en","voice":"af_heart","characters":32,"format":"m4a"}
curl -s "$URL" --output prompt.m4a # second step: download (or play) the linkWhat changes in practice: Polly is called through the AWS SDK with IAM credentials and a voice ID such as Joanna. Speakvora is a plain HTTPS call with an API key and voice IDs like af_heart, so it also works from places where AWS credentials are awkward, such as a Raspberry Pi, a kiosk image or a CI job. The text limit is 500 characters per request, and Speakvora returns JSON with a url that you fetch or play. Library links are stable public files; custom-voice links last about five minutes.
A safe migration order
- Stay on Polly while you remain inside the Standard free tier.
- When you leave it, move the fixed prompts to stored Speakvora audio first.
- Replace boto3 calls with one HTTPS request (code below).
- Store your own copy of audio you replay often, as you would with Polly.
Is it actually cheaper?
Polly's list prices are Standard $4, Neural $16, Generative $30 and Long-Form $100 per million characters. Speakvora's catalog voices are $0.70, about 82% below Standard, 96% below Neural and 98% below Generative. Standard is an older quality tier than Neural, so the closest quality comparison is Neural ($16).
Full numbers, scenarios and sources: Speakvora vs Amazon Polly.
Test these before you commit
- Work out whether you stay inside Polly's free tier of 5 million Standard characters a month; if you do, do not move.
- Compare Neural-tier voices against Speakvora's on a sample of your text, because Standard is an older tier.
- Check any SSML you use, such as breaks or phonemes; map it before migrating.
- Decide how you will store audio you replay; Polly states caching is free, and Speakvora pack audio is licensed for play inside your own apps.
Speakvora is in early access with billing is live; questions to support@eyeguide.ai. Latency: see the speed page.