YourTTS: Two Different Jobs, One Model
- What Is YourTTS?
- Built on VITS, With Real Modifications
- Two Distinct Capabilities
- A Genuine Contribution to Low-Resource Languages
- Fine-Tuning for Difficult Voices
- Getting Started with YourTTS
- Tips for Better Results
- YourTTS vs. Related Coqui and Voice-Transfer Models
- Frequently Asked Questions
- Give an Existing Recording a New Voice
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Generating new speech in an unseen voice and converting an existing recording into a different speaker's voice sound like related problems, but most TTS systems are built to do only one. YourTTS, developed by researchers at Coqui, was built to do both — zero-shot multi-speaker text-to-speech and zero-shot voice conversion, from the same underlying model — and in doing so became the first published approach to bring a genuinely multilingual method to zero-shot multi-speaker TTS.
What Is YourTTS?
YourTTS is a research model developed by Edresson Casanova and colleagues at Coqui, built on top of the VITS architecture — an end-to-end model that combines a variational autoencoder, adversarial training, and normalizing flows to go directly from text to a waveform without separate acoustic and vocoder stages. YourTTS layers several specific modifications onto that VITS foundation, tuned specifically for zero-shot, multilingual, multi-speaker generation.
Built on VITS, With Real Modifications
Two design choices stand out as genuine departures from prior work rather than incremental tuning. First, YourTTS uses raw text as input instead of phonemes — earlier systems typically relied on phonemization, which works well for languages with mature, high-quality grapheme-to-phoneme tools, but performs unevenly for languages that lack them. Working from raw text directly sidesteps that dependency, which matters specifically for extending the approach to lower-resource languages. Second, a dedicated speaker encoder extracts a speaker embedding from each training utterance to condition the model, trained alongside a Speaker Consistency Loss that compares embeddings extracted from generated audio against the original speaker's embedding using cosine distance — directly optimizing for the model actually sounding like the target speaker, not just producing intelligible speech.
Two Distinct Capabilities
Zero-shot multi-speaker TTS generates new speech in a voice the model has never been trained on, using only a reference embedding computed from a short sample. Zero-shot voice conversion is a related but genuinely different task: given an existing audio recording, transform it into a target speaker's voice while preserving the original content — you're converting what was already said, not generating new text. YourTTS achieved state-of-the-art results on the first task and results comparable to the state of the art on the second, both measured on the VCTK benchmark, which is what makes it a genuinely dual-purpose system rather than a TTS model with a voice conversion feature bolted on.
A Genuine Contribution to Low-Resource Languages
YourTTS demonstrated something specifically useful beyond its headline benchmark numbers: promising zero-shot results in a target language using just a single speaker's data during training. The published training setup reflects this directly — more than 1,000 English speakers were used, but only 5 French speakers and a single Portuguese speaker, with a language batch balancer deliberately allocating a larger share of training batches to the underrepresented languages to compensate. That a usable multilingual, multi-speaker system could emerge from such a lopsided dataset is the model's specific evidence that this approach could open zero-shot TTS and voice conversion to languages that don't have the deep speaker-diverse datasets that English does.
Fine-Tuning for Difficult Voices
For speakers whose voice or recording characteristics differ significantly from anything in the training data, YourTTS supports fine-tuning with less than a minute of additional speech, reported to achieve state-of-the-art voice similarity results even in these harder cases. This matters in practice for source material that's unusually accented, recorded in unusual conditions, or otherwise not well represented by the base model's zero-shot capability.
Getting Started with YourTTS
The most direct path is through the Coqui TTS library's command-line interface, specifying the YourTTS model along with a target speaker reference file, an optional content reference file for voice conversion, and a language code — the library handles model loading and inference from there. Given the broken links on the original research repository, pulling the model through Coqui TTS (or its actively maintained community forks) rather than the original paper's GitHub page is the more dependable route to a working checkpoint today.
Tips for Better Results
- Be clear about which task you're actually doing. Zero-shot TTS (new text, new voice) and voice conversion (existing audio, new voice) use YourTTS differently — decide which one your project needs before setting up your pipeline, since mixing up the workflow is a common source of confusion.
- Fine-tune when your target voice is unusual. If your source speaker has distinctive accent or recording characteristics not well represented in typical training data, budget for the documented sub-one-minute fine-tuning step rather than expecting zero-shot alone to nail it.
- Use the community-maintained Coqui TTS fork rather than the original repository. Given the broken links on Edresson's original GitHub page following Coqui's 2024 shutdown, the actively maintained library fork is the more reliable source for a working model.
- Respect consent when cloning or converting a real person's voice. Given documented research showing YourTTS's cloning capability can be used to spoof voice-based authentication, treat any real-world deployment involving a specific person's voice as requiring their explicit permission.
YourTTS vs. Related Coqui and Voice-Transfer Models
| YourTTS | XTTS-v2 | VoiceCraft | |
|---|---|---|---|
| Zero-shot multi-speaker TTS | Yes, state-of-the-art at release | Yes | Yes |
| Zero-shot voice conversion | Yes, comparable to SOTA at release | Not the primary focus | No (editing focus instead) |
| Base architecture | VITS + speaker encoder + Speaker Consistency Loss | GPT-2-style autoregressive + VQ-VAE | Autoregressive with causal masking + delayed stacking |
| Multilingual approach | First of its kind in this research space | 17 languages | English-focused |
| Text input | Raw text (no phonemizer dependency) | Standard text processing | Standard text processing |
| Current maintenance | Via Coqui TTS community forks | Via Coqui TTS community forks | Actively maintained by original authors |
YourTTS's specific place in this lineage is as the earlier Coqui research that established multilingual zero-shot multi-speaker TTS as a viable approach — work that the later XTTS models built further on, while YourTTS itself remains distinctive for treating voice conversion as a first-class capability rather than an afterthought.
Frequently Asked Questions
What's the difference between YourTTS's TTS mode and its voice conversion mode?
TTS mode generates new speech from text in a target voice; voice conversion takes an existing audio recording and transforms it into a target speaker's voice while preserving the original spoken content — two related but distinct tasks handled by the same model.
Is YourTTS still available to use?
Yes, primarily through the Coqui TTS library and its actively maintained community forks. The original research repository's model download links are broken following Coqui AI's 2024 shutdown and the wiping of its original hosting server.
How much data does YourTTS need to clone a voice?
Zero-shot cloning works from a short reference sample with no fine-tuning required; for voices that differ significantly from training data, fine-tuning with less than one minute of additional speech is documented to achieve notably better similarity.
Who developed YourTTS?
YourTTS was developed by Edresson Casanova and collaborators at Coqui, published at ICML 2022.
Is YourTTS free to use?
It's available through the open-source Coqui TTS library; check the current license terms of whichever specific package or fork you're using, since licensing details can vary across the community-maintained forks that emerged after Coqui AI's shutdown.
Give an Existing Recording a New Voice
For a browser-based workflow with a similar goal, use Echora’s Voice Changer. Upload an existing speech recording, choose a built-in target voice or an authorized reference sample, and generate transformed audio while keeping the spoken content and performance.