CosyVoice2: Streaming Low-Latency Multilingual TTS Model
CosyVoice2 streams speech with 150ms first-packet latency, cuts pronunciation errors 30-50% vs CosyVoice, and adds instructed emotion and accent control.
Guides to open-source and popular AI voice models, and how to try comparable capabilities on Echora.
CosyVoice2 streams speech with 150ms first-packet latency, cuts pronunciation errors 30-50% vs CosyVoice, and adds instructed emotion and accent control.
CosyVoice is Alibaba FunAudioLLM open-source TTS using supervised semantic tokens for zero-shot voice cloning across 5 languages, at just 300M parameters.
Kokoro-82M is an Apache-2.0 TTS model with 54 voices across 8 languages that runs fast on CPU with no GPU required. Free, open-source, easy to self-host.
Chatterbox Multilingual clones a voice from 5-10s of audio and speaks it in 23 languages, keeping accent and timbre consistent. Open-source, MIT-licensed.
Chatterbox Turbo distills TTS generation from 10 steps to 1, cutting latency to under 150ms for real-time voice agents. Fast, open-source, MIT-licensed.
Chatterbox is Resemble AI's open-source TTS model with zero-shot voice cloning, emotion control, and built-in watermarking. See how it compares with ElevenLabs.
Dia2 is Nari Labs' streaming upgrade to Dia TTS: real-time multi-speaker dialogue with consistent per-speaker voices as text arrives. See how it works.
Nari Labs' Dia 1.6B is an open-source speech model generating multi-speaker dialogue with real emotion, laughter, and sighs in one pass. Try it free online.