Step-Audio: StepFun Unified Speech Understanding + TTS
Step-Audio is a StepFun open-source model that unifies speech recognition, dialogue, and voice generation in one system, skipping the ASR-to-TTS handoff.
Guides to open-source and popular AI voice models, and how to try comparable capabilities on Echora.
Step-Audio is a StepFun open-source model that unifies speech recognition, dialogue, and voice generation in one system, skipping the ASR-to-TTS handoff.
Zonos is a Zyphra open-weight TTS model trained on 200k+ hours of speech, cloning a voice in seconds with fine control over emotion, pitch, and speaking rate.
Orpheus TTS is a Canopy Labs Llama-3B speech-LLM with 200ms streaming latency, emotion tags, and zero-shot cloning, deployable through a FastAPI server.
OuteTTS treats speech as pure language modeling, so it runs natively on llama.cpp via GGUF, cloning a voice from 10 seconds across 23 languages total.
GLM-TTS is Zhipu AI open-source TTS combining an LLM with flow matching, cloning a voice from 3 seconds and using RL to boost emotion and pronunciation.
Parler-TTS is a Hugging Face open-source TTS model that controls voice, pitch, and speaking style through a plain-language description, no audio needed.
ChatTTS is a 2noise TTS model optimized for chatbot and assistant dialogue in Chinese and English, generating natural laughter, pauses, and interjections.
MeloTTS is MyShell open-source multilingual TTS with real-time CPU inference and native Chinese-English code-switching within a single spoken sentence.
Piper TTS is a 15M-parameter neural voice model that runs in real time on CPU alone, built for Raspberry Pi and other edge devices with no GPU needed.
NaturalSpeech 3 splits speech into separate content, prosody, timbre, and detail subspaces, letting Microsoft TTS model control each one independently.
NaturalSpeech 2 uses latent diffusion on continuous audio codes, letting Microsoft TTS model generate zero-shot singing from just a spoken voice prompt.
NaturalSpeech is Microsoft VAE-based TTS model that formally defined human-level quality, then passed its own statistical test on the LJSpeech dataset.