Silero TTS: Lightweight Pretrained Text-to-Speech Models
Silero TTS is a lightweight, MIT-licensed pretrained model loadable in one line via PyTorch Hub, with strong Russian, CIS, and Indic language coverage.
Guides to open-source and popular AI voice models, and how to try comparable capabilities on Echora.
Silero TTS is a lightweight, MIT-licensed pretrained model loadable in one line via PyTorch Hub, with strong Russian, CIS, and Indic language coverage.
VITS is the 2021 conditional VAE and flow-based TTS model that proved single-stage generation could match two-stage systems, spawning countless forks.
MockingBird clones a voice from just 5 seconds of audio for real-time Mandarin speech synthesis, and remains one of the most popular Chinese TTS repos.
EmotiVoice is a NetEase Youdao open-source TTS engine with 2,000+ voices in English and Chinese, letting you set the mood through a simple text prompt.
YourTTS is a Coqui research model built on VITS, the first multilingual zero-shot multi-speaker TTS system that also performs zero-shot voice conversion.
WhisperSpeech reverses OpenAI's Whisper speech-recognition architecture to generate speech from text, using only properly licensed, commercially safe data.
VoiceCraft edits existing speech and clones voices from a few seconds of audio, with edited clips preferred over the original recording 48% of the time.
Voicebox is a Meta AI research model that infills speech from text and audio context, using flow matching to power zero-shot TTS, editing, and denoising.
MetaVoice is an Apache 2.0 TTS model built for emotional English speech, cloning American and British voices from just 30 seconds of reference audio.
VibeVoice is Microsoft open-source TTS generating 90 minutes of 4-speaker podcast audio, with a companion VibeVoice-ASR model and 4-bit quantized variants.
Sesame CSM 1B generates speech from interleaved text and audio history, matching a conversation's actual pace and emotion instead of reading flat text.
Llasa extends LLaMA 1B, 3B, and 8B models with XCodec2 speech tokens, trained on 250k hours of Chinese-English audio for voice cloning and text-to-speech.