EchoraEchora

AI Models

Guides to open-source and popular AI voice models, and how to try comparable capabilities on Echora.

7 min read

Step-Audio: StepFun Unified Speech Understanding + TTS

Step-Audio is a StepFun open-source model that unifies speech recognition, dialogue, and voice generation in one system, skipping the ASR-to-TTS handoff.

Step-Audiospeech interactionspeech understandingtext to speechvoice agentsopen source
Read more
7 min read

Parler-TTS: Hugging Face Text-Described Voice Model

Parler-TTS is a Hugging Face open-source TTS model that controls voice, pitch, and speaking style through a plain-language description, no audio needed.

Parler-TTStext to speechHugging Facenatural-language voice controlspeech synthesisopen source
Read more
7 min read

MeloTTS: MyShell Multilingual TTS That Runs on CPU

MeloTTS is MyShell open-source multilingual TTS with real-time CPU inference and native Chinese-English code-switching within a single spoken sentence.

MeloTTSmultilingual TTSCPU TTScode-switchingspeech synthesisopen source
Read more
6 min read

NaturalSpeech 3: Microsoft Factorized Zero-Shot TTS

NaturalSpeech 3 splits speech into separate content, prosody, timbre, and detail subspaces, letting Microsoft TTS model control each one independently.

NaturalSpeech 3TTSFACodecfactorized diffusionzero-shot speech synthesisMicrosoft Research
Read more
7 min read

NaturalSpeech 2: Microsoft Zero-Shot Singing TTS Model

NaturalSpeech 2 uses latent diffusion on continuous audio codes, letting Microsoft TTS model generate zero-shot singing from just a spoken voice prompt.

TTSNaturalSpeech 2zero-shot singing synthesisvoice cloninglatent diffusionMicrosoft Research
Read more