Amazon Polly: AWS Cloud Text-to-Speech, Four Engines
Amazon Polly is AWS's cloud TTS service offering four distinct voice engines, from cheap Standard voices to premium Long-Form narration, tied into AWS.
Guides to open-source and popular AI voice models, and how to try comparable capabilities on Echora.
Amazon Polly is AWS's cloud TTS service offering four distinct voice engines, from cheap Standard voices to premium Long-Form narration, tied into AWS.
Azure AI Speech is Microsoft's official cloud TTS service, offering hundreds of neural voices, custom voice training, and full SSML control at real scale.
Edge TTS is a free, open-source tool that taps into the same online voices behind Microsoft Edge's Read Aloud feature, with no API key or account needed.
SeamlessM4T is a Meta open-source model unifying speech recognition, translation, and speech synthesis into a single system across up to 100 languages.
MaskGCT is an Amphion non-autoregressive TTS model that skips alignment and duration prediction entirely, cloning a voice from just 5 seconds of audio.
FireRedTTS is a Xiaohongshu foundation TTS framework offering zero-shot voice cloning for dubbing and instruction-tuned casual speech for AI chatbots.
Spark-TTS uses BiCodec, a single-stream speech codec built on Qwen2.5, for zero-shot voice cloning and fully synthetic voice creation in one open model.
NeuTTS Air is a Neuphonic on-device TTS model under 1B parameters that clones a voice from just 3 seconds of audio, running on phones and Raspberry Pi.
Kyutai Pocket TTS is a 100M-parameter, MIT-licensed model that runs real-time speech synthesis and zero-shot voice cloning on CPU alone, no GPU required.
KaniTTS pairs a compact LFM2 language model with an efficient neural audio codec, running real-time speech synthesis on as little as 3GB of GPU VRAM total.
VoxCPM is an OpenBMB tokenizer-free TTS model for Chinese and English, hitting a 0.17 real-time factor on a consumer GPU, with automatic context-aware prosody.
Voxtral TTS is a 4B open-weight model from Mistral AI that clones a voice from 2-3 seconds of audio, benchmarked directly against ElevenLabs on quality.