MeloTTS: The Model That Doesn't Stumble When a Sentence Switches Languages
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Say a sentence with an English brand name dropped into the middle of otherwise fluent Chinese, and most TTS models either mangle the English word or force an awkward pause around it. MeloTTS, developed jointly by MIT and MyShell.ai, was specifically built to handle exactly that situation — its Chinese voice model reads mixed Chinese-English text naturally, switching pronunciation mid-sentence the way an actual bilingual speaker would, without treating it as an edge case to work around.
What Is MeloTTS?
MeloTTS is an open-source, MIT-licensed multilingual text-to-speech library built on the VITS and VITS2 architecture, with BERT-based linguistic features layered in to improve how naturally it captures prosody and phrasing. It covers English, Spanish, French, Chinese, Japanese, and Korean, and — notably for a project of this size — ships five distinct English accent variants: American, British, Indian, Australian, and a general default, all from the same underlying project rather than requiring separate models per accent.
The BERT-based linguistic feature extraction is worth understanding as more than a technical footnote: rather than relying purely on a rule-based phonemizer to convert text into pronunciation units, MeloTTS pulls in contextual linguistic understanding from a BERT-style model, which is part of why its prosody and phrasing tend to sound more natural than TTS systems relying on simpler grapheme-to-phoneme conversion alone.
Native Chinese-English Code-Switching
This is MeloTTS's most distinctive capability, and it solves a genuinely common, genuinely annoying problem. Bilingual or code-mixed content — a Chinese sentence referencing an English product name, a tech term, or a quoted phrase — trips up most TTS systems, which typically expect input in a single consistent language and either mispronounce the foreign segment or require you to split it into separate generation calls. MeloTTS's Chinese speaker model handles mixed Chinese-English text directly, switching pronunciation naturally within a single sentence and a single generation call, without manual language-tagging or text splitting on your end.
CPU Real-Time Performance
Like a couple of other lightweight models worth knowing about if you've compared options in this space, MeloTTS is built to run comfortably on CPU alone — the project's own documentation is explicit that CPU is sufficient for real-time inference, with GPU support available (cuda, cuda:0, or mps on Apple Silicon) as an optional accelerator rather than a requirement. This makes it a genuine alternative for teams that want multilingual coverage without provisioning GPU infrastructure, though it's worth noting MeloTTS prioritizes broader language and accent coverage over the absolute minimal footprint of something like Piper, which trades multilingual breadth for extreme lightness on constrained edge hardware.
Getting Started with MeloTTS
- Clone the repository and install dependencies.
git clonethemyshell-ai/MeloTTSproject from GitHub and follow the documented installation steps before running any inference code. - Use the simple Python API for basic generation.
from melo.api import TTS, then instantiate with a language code (EN,ES,FR,ZH,JP,KR, or the newerEN_NEWESTvariant) and a device setting (cpu,cuda,cuda:0,mps, orautoto let it choose automatically). - Select your specific English accent if relevant. Since MeloTTS ships American, British, Indian, and Australian English variants distinctly, pick the one matching your target audience rather than defaulting to the general English model.
- Adjust the speed parameter to fit your content. Generation speed is directly adjustable in the API call, useful for matching narration pace to video timing or personal preference without needing a separate post-processing step.
- Use the community-contributed WebUI or CLI for quicker testing. Both were added by a community contributor on top of the core library, giving you a way to test voices and languages without writing a Python script for every experiment.
- Check for regional and community-extended variants if your language isn't natively covered. Community forks — a dedicated Vietnamese variant is one example — have extended MeloTTS's approach to languages with phonological features (like Vietnamese's six tones) that the original phonemizer wasn't built to handle.
Tips for Better Results
- Lean into the mixed-language capability deliberately. If your content naturally includes bilingual phrasing — product names, technical terms, quoted foreign phrases — feeding it through as a single mixed-language sentence rather than splitting it into separate calls per language is exactly the workflow MeloTTS's Chinese model was built to handle.
- Pick the right English accent for your audience up front. Choosing British English for a UK-facing project or Indian English for that specific market gives a noticeably better fit than defaulting to the general American-leaning model and hoping it reads as neutral.
- Test CPU inference on your actual target hardware before assuming you need a GPU. Given the project's explicit real-time-on-CPU design, it's worth confirming performance on your intended deployment machine before provisioning GPU resources you may not need.
- Check community forks first if you need a language outside the core six. Rather than trying to force an unsupported language through MeloTTS's default phonemizer, search for an existing community extension — several exist specifically because the default text-processing pipeline doesn't handle every language's phonology correctly out of the box.
- Use the newer
EN_NEWESTvariant for English content if quality is the priority. Iterative English-specific releases (v2, v3) suggest ongoing refinement specifically for that language — worth testing against the originalENmodel rather than assuming they're interchangeable.
MeloTTS vs. Other CPU-Friendly Multilingual Models
| MeloTTS | Piper | Kokoro-82M | |
|---|---|---|---|
| Backing organization | MIT + MyShell.ai | Rhasspy / Open Home Foundation | Independent (hexgrad) |
| Core languages | 6 (plus 5 English accent variants) | 35+ | 8 |
| Architecture | VITS/VITS2 + BERT linguistic features | VITS + ONNX + espeak-ng | StyleTTS 2 + ISTFTNet |
| Mixed-language (code-switching) | Yes, native to Chinese model | Not a dedicated feature | Not a dedicated feature |
| CPU real-time | Yes | Yes, including Raspberry Pi | Yes |
| Voice cloning | No, fixed voice roster | No, fixed voice roster | No, fixed voice roster |
| License | MIT | MIT (original) / GPL-3.0 (current fork) | Apache 2.0 |
MeloTTS's specific niche among the CPU-friendly, non-cloning TTS options is depth over breadth in a particular direction: fewer total languages than Piper, but genuine linguistic sophistication — BERT-based prosody and native code-switching — in the languages it does support, especially for bilingual Chinese-English content.
Frequently Asked Questions
Who developed MeloTTS?
MeloTTS was developed jointly by MIT and MyShell.ai, with the project first released in 2023 and actively maintained on GitHub and Hugging Face since.
What makes MeloTTS good for Chinese-English content specifically?
Its Chinese speaker model natively handles mixed Chinese-English text, switching pronunciation correctly within a single sentence and generation call — a capability most TTS systems don't handle without manual language splitting.
Does MeloTTS need a GPU?
No. It's explicitly designed for real-time CPU inference, with GPU support available as an optional accelerator rather than a requirement.
Can MeloTTS clone a specific person's voice?
No. It works from a fixed set of speaker voices per language and accent rather than zero-shot voice cloning from a reference clip.
Is MeloTTS free for commercial use?
Yes. It's released under the MIT license, which permits both commercial and non-commercial use without royalties.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.