EmotiVoice: Set the Mood by Just Writing It Down
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most "emotional" TTS models give you a short menu of fixed emotion tags to pick from. EmotiVoice, open-sourced by NetEase Youdao, takes a more flexible approach: describe the mood you want in a style prompt, the same way you'd direct a voice actor, and the model performs it. That prompt-driven design — rather than a locked set of switches — is the model's defining trait, built on a bilingual foundation spanning English and Chinese with a genuinely large voice library to draw from.
What Is EmotiVoice?
EmotiVoice is an open-source, Apache 2.0-licensed text-to-speech engine released by NetEase Youdao in November 2023. It's built as a transformer-based system directly inspired by the PromptTTS research approach, and trained on the LibriTTS and Hi-Fi TTS datasets. The project positions itself explicitly as a multi-voice, prompt-controlled engine — both halves of that description matter equally to what makes it distinctive.
Functional Advantages of EmotiVoice
- Over 2,000 preset voices spanning English and Chinese, covering a genuinely wide range of ages, tones, and vocal characters rather than a handful of generic defaults.
- Prompt-controlled emotional delivery — happy, excited, sad, angry, and other moods are set through a style prompt rather than a rigid fixed-tag system.
- Additional style factors beyond emotion, including pitch, speed, and energy, independently adjustable alongside mood.
- Voice cloning from your own recordings, added shortly after the initial release, letting you extend the preset library with a custom voice.
- Multiple integration paths: an interactive web interface, a scripting interface for batch generation, and an OpenAI-compatible TTS API for dropping into existing tooling with minimal changes.
- Fully open and self-hostable, running offline once deployed, under a license with no commercial restrictions.
How Prompt-Based Style Control Actually Works
Built on the PromptTTS approach, EmotiVoice's current implementation uses four style factors to shape delivery: pitch, speed, energy, and emotion — notably not gender, which the project's own documentation specifically calls out as excluded from the current style-control set, though it notes this could be added without major difficulty if there's demand. In practice, this means a voice's fundamental gender characteristic comes from the preset voice you choose, while everything about how that voice delivers a line — its emotional coloring, pacing, and intensity — is what the prompt actually shapes. This split is worth understanding before you start writing prompts: pick your voice first for identity, then write your prompt to direct the performance.
Getting Started with EmotiVoice
The fastest path to trying EmotiVoice is its Docker image, which requires a machine with an NVIDIA GPU and the NVIDIA Container Toolkit configured — running the provided image launches an interactive web demo you can use directly in the browser. If you'd rather set things up manually, a local installation is documented using a Python 3.8 environment via conda, giving you more control over configuration and dependencies. For programmatic use, the community-contributed OpenAI-compatible API server can be installed with a short list of Python packages and run alongside the core engine, letting you call EmotiVoice from any client already built for OpenAI's TTS API format with minimal code changes.
Tips for Better Results
- Write style prompts the way you'd direct a performance, not just label an emotion. Since the system is prompt-driven rather than tag-based, a more descriptive prompt covering intensity and pacing alongside the core emotion tends to produce more precise results than a single-word label.
- Choose your preset voice before writing your emotional prompt. Since gender and core vocal identity come from voice selection rather than the style prompt, settle on the right voice first, then iterate on the prompt to get the delivery you want.
- Use the batch scripting interface for volume work. If you're generating more than a handful of clips, the scripting interface avoids the overhead of manually running the web interface for each one.
- Reach for voice cloning specifically when none of the 2,000+ presets fit. Given how large the existing voice library already is, it's worth browsing presets first — cloning is there for the cases where you need a specific real voice, not as the default starting point.
- Use the OpenAI-compatible API if you're replacing a commercial TTS vendor. Since it mirrors that API's request format, migrating existing integration code is typically a smaller change than adapting to a fully custom interface.
EmotiVoice vs. Other Expressive Bilingual Models
| EmotiVoice | ChatTTS | Chatterbox | |
|---|---|---|---|
| Backing organization | NetEase Youdao | 2noise | Resemble AI |
| Languages | English, Chinese | Chinese, English | English (Multilingual: 23) |
| Emotion control method | Free-form style prompt | Predicted natively during generation | Explicit exaggeration parameter |
| Preset voice count | 2,000+ | Multi-speaker via Gaussian sampling | Zero-shot cloning focus |
| Voice cloning | Yes, from personal recordings | Not the primary workflow | Yes, zero-shot |
| License | Apache 2.0 | AGPLv3+ (code) / CC BY-NC 4.0 (weights) | MIT |
EmotiVoice's specific niche is the combination of prompt-based flexibility with sheer preset-voice scale — where ChatTTS bakes conversational prosody directly into generation and Chatterbox exposes a single exaggeration dial, EmotiVoice lets you describe delivery in your own words across a genuinely large voice catalog.
Frequently Asked Questions
How do I control emotion in EmotiVoice?
Through a style prompt written in plain language, rather than selecting from a fixed set of emotion tags — the model interprets the prompt to shape delivery alongside separately adjustable pitch, speed, and energy settings.
Can EmotiVoice control a voice's gender through prompts?
Not currently. Gender comes from the preset voice you select rather than the style prompt, since the project's current style-control implementation deliberately excludes gender as a factor, though this is noted as something that could be added in the future.
Does EmotiVoice support voice cloning?
Yes, using your own recordings, in addition to its library of more than 2,000 preset voices.
What languages does EmotiVoice support?
English and Chinese, with additional language support (including Japanese and Korean) noted on the project's roadmap.
Is EmotiVoice free for commercial use?
Yes. It's released under the Apache 2.0 license, which permits commercial use without royalties or separate licensing agreements.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.