EchoraEchora
Back to Models

ChatTTS: Built for Talking, Not Reading Aloud

September 4, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Most TTS models are optimized to read a passage of text clearly and pleasantly — which is exactly the wrong target for a chatbot or voice assistant that's supposed to sound like it's actually talking to you. Real conversation is full of small pauses, an occasional laugh, an "um" that isn't in the script. ChatTTS, developed by the 2noise team, was built specifically around that distinction — optimized not for narration, but for the messier, more natural texture of actual dialogue.

What Is ChatTTS?

ChatTTS is a 300-million-parameter generative speech model purpose-built for conversational scenarios — LLM assistant dialogue, chatbot voices, and interactive voice applications — rather than general narration or audiobook-style reading. It supports both Chinese and English, and the full model was trained on more than 100,000 hours of bilingual conversational data, with the specific goal of capturing how people actually sound when they talk, not just how text sounds when it's read.

Fine-Grained Prosodic Control: Laughter, Pauses, Interjections

This is ChatTTS's defining capability, and it's built into the model's core generation process rather than bolted on as a post-processing effect. The model can predict and control fine-grained prosodic features directly — laughter, natural pauses, and conversational interjections — as part of generating the speech itself, rather than requiring you to manually insert sound effect tags the way some other expressive models do. The result, according to the project's own comparisons, is prosody that surpasses most open-source TTS alternatives specifically in how natural and conversational the output feels, rather than how clearly it reads.

Multi-Speaker Generation via Gaussian Sampling

ChatTTS supports multiple speakers for genuinely interactive, multi-party conversations — useful for scenarios like simulated dialogues or multi-character assistant interactions. Its approach to speaker variation is worth understanding: rather than requiring a reference audio clip for each voice, speakers can be sampled directly from a Gaussian distribution within the model's own inference process, giving you a straightforward way to generate a variety of distinct voices without needing a library of pre-recorded reference clips for each one.

An Honest Note on What's Actually Available

Two details are worth knowing before you build on ChatTTS, since they shape what you can realistically expect from a self-hosted setup.

The open-source checkpoint isn't the full model. ChatTTS's best-performing internal version was trained on the full 100,000+ hours of data with supervised fine-tuning (SFT) applied afterward. The version publicly released on Hugging Face is a 40,000-hour pretrained checkpoint without that SFT step — a genuinely capable base model, but not the same version behind the project's most polished demo comparisons.

Licensing is split between code and weights. ChatTTS's code is released under AGPLv3+, while the model weights themselves are released under a CC BY-NC 4.0 license — meaning the pretrained weights are restricted to non-commercial use. If commercial deployment is part of your plan, this is worth reviewing carefully before building a product around the public checkpoint.

Getting Started with ChatTTS

  1. Clone the repository. git clone https://github.com/2noise/ChatTTS.git gets you the code and setup instructions; installing via pip install chattts-fork is also documented, though it downloads close to a gigabyte of content on first install.
  2. Load the model and run a basic generation. A few lines of Python — import ChatTTS, chat = ChatTTS.Chat(), chat.load_models() — get you to your first generated clip; setting compile=True on model loading improves inference performance at the cost of a longer initial load.
  3. Watch the punctuation rules closely. ChatTTS's text input specifically expects English commas and periods as punctuation — other punctuation marks are treated as invalid input, which is a common source of unexpected generation failures for new users porting in text with different punctuation conventions.
  4. Sample speakers from the Gaussian distribution for variety. Rather than sourcing reference clips for each voice you want, use the model's built-in speaker sampling to generate a range of distinct voices directly.
  5. Test on modest hardware before assuming you need a powerful GPU. Community benchmarking on a LattePanda Sigma mini-PC reported generating a 22-second clip in about 39 seconds — a roughly 1:2 processing ratio that's been described as comparable to performance on an RTX 4090-equipped machine, suggesting the model is more forgiving of constrained hardware than its parameter count might suggest.
  6. Check current license terms before commercial use. Given the split AGPLv3+/CC BY-NC 4.0 licensing, confirm your specific use case is covered, particularly if the pretrained weights are involved in a paid product.

Tips for Better Results

  • Design your text specifically for conversation, not narration. ChatTTS is tuned for dialogue-style content — feeding it stiff, formally structured prose won't showcase what it's actually built for; conversational phrasing with natural pauses will get more out of its prosody modeling.
  • Sanitize punctuation before generation. Given the model's strict expectations around English commas and periods, running your input text through a quick cleanup pass to strip unsupported punctuation avoids a common, easily preventable failure point.
  • Use the Gaussian speaker sampling for quick multi-voice prototyping. It's a faster way to test multi-speaker dialogue scenarios than sourcing and preparing reference clips for each character upfront.
  • Set expectations around the public checkpoint's capability. Since the openly released model is the 40,000-hour pre-SFT version rather than the full 100,000+ hour fine-tuned model, calibrate your quality expectations accordingly rather than assuming you're getting the exact model shown in the project's best demo clips.
  • Budget realistic generation time on constrained hardware. The LattePanda Sigma benchmark is a useful data point if you're considering edge or mini-PC deployment — plan around roughly 2x real-time processing rather than assuming instant generation on non-GPU-class hardware.

ChatTTS vs. Other Expressive and Bilingual Models

ChatTTSBarkCosyVoice3
Primary design targetConversational dialogue, chatbots/assistantsGenerative text-to-audio (speech, music, effects)Multilingual cloning with dialects
LanguagesChinese, English13+9 + 18 Chinese dialects
Non-verbal soundsPredicted natively (laughter, pauses, interjections)Bracket tags ([laughter], [sighs])Not a dedicated feature
Multi-speakerYes, via Gaussian samplingPreset voicesZero-shot cloning
Parameters300MNot officially disclosed at this scale0.5B
Weights licenseCC BY-NC 4.0 (non-commercial)MITOpen, hosted API available

ChatTTS's specific niche is dialogue-first design for exactly two languages — rather than chasing broad multilingual coverage or maximum expressive range, it concentrates on making Chinese and English conversational speech sound like an actual conversation, which is the more relevant comparison point than raw language count for anyone building an assistant or chatbot voice.

Frequently Asked Questions

What makes ChatTTS different from a general-purpose TTS model?

It's specifically optimized for conversational, dialogue-style speech rather than narration — predicting natural pauses, laughter, and interjections directly as part of generation, which is tuned for chatbot and assistant use cases rather than reading text aloud clearly.

Is the publicly available ChatTTS model the full version?

Not exactly. The full internal model was trained on 100,000+ hours of data with supervised fine-tuning applied. The publicly released Hugging Face checkpoint is a 40,000-hour pretrained version without that fine-tuning step.

Is ChatTTS free for commercial use?

The code is under AGPLv3+, but the model weights are released under a CC BY-NC 4.0 license, which restricts commercial use — review current license terms carefully before any commercial deployment.

Why does ChatTTS reject some punctuation in my input text?

The model's text processing specifically expects English commas and periods; other punctuation marks are treated as invalid input, so cleaning up your text beforehand avoids common generation failures.

Does ChatTTS support voice cloning?

Not in the way dedicated cloning models do. It supports multi-speaker generation through Gaussian sampling of speaker characteristics rather than reproducing a specific reference voice from an audio clip.

Turn a Multi-Speaker Script into Audio

If you already have a finished multi-speaker script, use Echora’s Text to Dialogue. Add up to 12 dialogue blocks with 5,000 characters in total, assign a voice to each block, add supported delivery tags, and adjust stability or speaker boost. You can then preview and download the complete conversation.

Create Dialogue Audio →