GPT-SoVITS: 1-Minute Voice Cloning, Strong in Chinese
GPT-SoVITS clones a voice from just 1 minute of audio, or 5 seconds zero-shot, with especially strong Chinese and Japanese support and a full WebUI toolkit.
Guides to open-source and popular AI voice models, and how to try comparable capabilities on Echora.
GPT-SoVITS clones a voice from just 1 minute of audio, or 5 seconds zero-shot, with especially strong Chinese and Japanese support and a full WebUI toolkit.
VALL-E 2 adds repetition-aware sampling and grouped code modeling to VALL-E, reportedly reaching human parity in zero-shot voice cloning benchmark tests.
VALL-E X is Microsoft cross-lingual extension of VALL-E, cloning a voice from one language to speak fluently in another, with built-in accent control.
VALL-E is Microsoft neural codec language model that treats speech as discrete tokens like GPT treats text, cloning a voice from a 3-second audio prompt.
StyleTTS 2 uses diffusion to sample speaking style with no reference audio needed, reportedly surpassing human recordings in naturalness on the LJSpeech set.
StyleTTS separates speaking style from content using AdaIN and a dedicated style encoder, giving non-autoregressive TTS controllable prosody and delivery.
Bark is Suno open-source generative text-to-audio model that adds laughter, sighs, and music through simple bracket tags, released under an MIT license.
Tortoise TTS is an Apache 2.0 open-source multi-voice model combining an autoregressive transformer with diffusion for exceptional, if slow, audio quality.
XTTS-v2 is Coqui zero-shot voice cloning model spanning 17 languages from just a 6-second clip, still a benchmark despite Coqui shutting down in 2024.
E2 TTS is Microsoft embarrassingly simple zero-shot TTS using flow matching with no duration model or aligner, and the direct research ancestor of F5-TTS.
F5-TTS is a non-autoregressive Diffusion Transformer TTS model using flow matching, cloning voices in seconds with a fast 0.15 real-time inference factor.
Fish Speech is Fish Audio expressive open-source TTS with zero-shot cloning and LoRA fine-tuning, plus emotion tags like [whisper] and [angry] built in.