Amazon Polly: Four Price Tiers, So You Only Pay for the Quality You Actually Need
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most TTS APIs give you one voice quality tier and one price point, whether you're generating a quick IVR prompt or a full audiobook chapter. Amazon Polly, AWS's text-to-speech service, takes a deliberately different approach: four separate voice engines, each targeting a different quality-and-cost trade-off, so a high-volume notification system isn't paying premium rates for expressiveness it doesn't need, and a narration-heavy project isn't stuck with robotic-sounding output to save money.
What Is Amazon Polly?
Amazon Polly is AWS's cloud-based text-to-speech service, using deep learning to convert written text into natural-sounding speech across dozens of languages. It's built as a core AWS service, meaning it integrates natively with the rest of the AWS ecosystem — S3 for audio storage, CloudWatch for monitoring, IAM for access control — rather than functioning as a standalone product bolted onto AWS as an afterthought.
Functional Advantages of Amazon Polly
- Four distinct voice engines at four different price points, letting you deliberately match cost to the quality your specific content actually needs rather than paying one flat premium rate for everything.
- Broad language coverage, spanning dozens of languages and regional voice variants.
- Deep native AWS integration, working directly with S3, CloudWatch, IAM, and Lambda for teams already building on AWS infrastructure.
- A genuinely generous 12-month free tier, with a meaningful character allowance across all four engines rather than just the cheapest one.
- A stated privacy commitment: Amazon Polly does not retain the content of your text submissions.
- Comprehensive SSML support, giving direct control over pronunciation, pacing, and emphasis across all engine tiers.
The Four Engines: Picking the Right Tier for Your Content
This tiered structure is Polly's defining characteristic, and picking correctly matters more than defaulting to the priciest option. Standard voices are the original, most affordable tier — solid for basic notifications, IVR prompts, and other high-volume, low-nuance use cases where clarity matters more than natural expressiveness. Neural voices, built on deep learning, are the tier most commonly recommended for production use — smoother intonation, more natural pacing, and meaningfully better handling of complex sentences, at a real but justified premium over Standard. Generative voices, added in 2024 and significantly expanded in March 2026, use a billion-parameter transformer model to push further into expressive, nuanced delivery — the tier to reach for when content genuinely benefits from more natural-sounding performance than Neural provides. Long-Form voices are purpose-built for a different problem entirely: extended listening. Rather than optimizing for a single line or paragraph, Long-Form is tuned specifically to reduce listener fatigue across genuinely long content like audiobooks or full articles — the highest-cost tier, reserved for content where that specific listening-duration optimization actually pays off.
What Polly Doesn't Do
Worth knowing honestly before you build around it: Amazon Polly doesn't offer voice cloning as a core capability, unlike some competing cloud providers that have added restricted cloning features. Independent reviews also note a real gap in creating highly distinct, branded voices without reaching for more advanced (and separately priced) customization options, and that for genuinely dramatic, highly emotional delivery, dedicated expressive-TTS competitors can still outperform Polly's most expressive tier. If your project's core requirement is reproducing a specific real voice or maximizing raw emotional range, it's worth evaluating Polly's Generative tier directly against those alternatives before committing, rather than assuming the newest engine closes every gap.
Getting Started with Amazon Polly
- Set up an AWS account and IAM credentials. Since Polly is accessed through standard AWS authentication, you'll need an access key and secret key with the appropriate Polly permissions attached to your IAM user or role before making any API calls.
- Install the AWS SDK for your language. For Python,
boto3is the standard client library —pip install boto3gets you started, with equivalent SDKs available for JavaScript, Java, and other supported languages. - Call
synthesize_speechwith your text and chosen engine. A basic Python call looks likepolly_client.synthesize_speech(Text="your text", OutputFormat="mp3", VoiceId="Joanna", Engine="neural")— theEngineparameter is where you select Standard, Neural, Generative, or Long-Form specifically. - Check which voices support which engine before assuming full compatibility. Not every voice is available across all four engines — confirm your chosen
VoiceIdactually supports the engine tier you want before building a pipeline around that combination. - Use SSML tags directly in your text input for finer control. Wrapping input in SSML markup gives you pronunciation, pacing, and emphasis control across any of the four engines.
- Monitor usage through CloudWatch and budget for AWS-adjacent costs. Beyond the per-character engine pricing, factor in data transfer fees and S3 storage costs if you're saving generated audio files, since these are billed separately from the core Polly usage.
Tips for Better Results
- Default to Neural for most production content, not Standard. Given how meaningfully more natural Neural sounds for a moderate cost increase over Standard, it's the more sensible default for anything customer-facing rather than treating Standard as the safe starting point.
- Reserve Generative and Long-Form for content that specifically benefits from them. Generative's expressiveness and Long-Form's fatigue-reduction tuning are both real advantages, but at meaningfully higher per-character cost — match the tier to content that actually needs that specific quality rather than defaulting to the top tier everywhere.
- Test your specific voice-and-engine combination before committing. Since not all voices support all four engines, validate availability for your target language and voice early rather than discovering a mismatch after building your pipeline around it.
- Check current pricing directly before budgeting at scale. Amazon has adjusted Polly's engine pricing over time, and quoted per-character rates can vary by source — confirm the current numbers on AWS's own pricing page rather than relying on a cited figure that may be outdated.
- Look elsewhere if voice cloning is a core requirement. Since Polly doesn't offer this as a built-in capability, a provider that specifically supports cloning (with appropriate consent and safeguards) is the more direct fit for that specific need.
Amazon Polly vs. Azure AI Speech and Edge TTS
| Amazon Polly | Azure AI Speech | Edge TTS | |
|---|---|---|---|
| Pricing structure | Four tiers ($4–$100+ per 1M characters) | Two main tiers (Neural, Neural HD) | Free |
| Voice cloning | Not a core feature | Yes, via restricted Personal Voice | Not applicable |
| Ecosystem integration | Deep AWS integration (S3, CloudWatch, IAM) | Deep Azure/Foundry integration | None — standalone tool |
| Free tier | Yes, 12 months, generous allowance | Yes, monthly allowance | Unlimited, free by nature |
| Official, licensed API | Yes | Yes | No — unofficial community tool |
| Best fit | Teams already on AWS, tiered cost control | Teams needing SSML depth and custom voice training | Prototyping and personal projects |
Amazon Polly's specific niche is deliberate cost-tiering combined with deep AWS-native integration — the right default for any team already building on AWS infrastructure that wants explicit control over the price-to-quality trade-off rather than one fixed rate for every use case.
Frequently Asked Questions
What's the difference between Polly's four voice engines?
Standard is the most affordable, best suited to basic notifications and IVR. Neural offers noticeably more natural intonation and pacing at a moderate premium, recommended for most production content. Generative uses a billion-parameter transformer model for more expressive, nuanced delivery. Long-Form is specifically tuned to reduce listener fatigue across extended content like audiobooks.
Does Amazon Polly support voice cloning?
Not as a core, generally available feature. If cloning a specific voice is a central requirement, a provider that offers this capability directly (with appropriate consent safeguards) is a better fit.
Is Amazon Polly free to use?
It offers a 12-month free tier for new AWS customers, with character allowances varying by engine — Standard's allowance is considerably larger than Generative's or Long-Form's. Beyond that free tier, it's a pay-as-you-go service billed per character.
Does Amazon Polly retain the text I submit?
No — Amazon states directly that Polly does not retain the content of text submissions.
Which engine should I use for a customer-facing application?
Neural is generally the recommended starting point for customer-facing content, balancing natural-sounding quality against cost; reserve Generative specifically for content where its added expressiveness is worth the higher price.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora's Text to Speech. Turn up to 5,000 characters into single-speaker audio, preview the result, and download the finished MP3.