EchoraEchora
Back to Models

XTTS-v2: The Voice Cloning Benchmark That Outlived Its Own Company

August 30, 2026
•
8 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Plenty of open-source models get compared to XTTS-v2 before they get compared to anything else. That's a strange position for a model to hold, considering the company that built it, Coqui AI, shut down entirely in early 2024. XTTS-v2 is the model that made that reputation stick anyway — a multilingual zero-shot voice cloning system good enough that its own creator's disappearance barely dented its relevance.

What Is XTTS-v2?

XTTS-v2 is an open-source multilingual text-to-speech and voice cloning model developed by Coqui AI, released in 2023 as a refinement of the original XTTS architecture. It clones a speaker's voice from as little as 6 seconds of reference audio, across 17 languages: Arabic, Brazilian Portuguese, Mandarin Chinese, Czech, Dutch, English, French, German, Italian, Polish, Russian, Spanish, Turkish, Japanese, Korean, Hungarian, and Hindi.

Architecturally, XTTS-v2 is GPT-based rather than diffusion-based — a meaningful distinction from newer flow-matching models in the same space. A VQ-VAE first converts ground-truth audio into discrete acoustic codes, which serve as training targets. A GPT-2-style language model then predicts these audio tokens conditioned on the input text and a speaker latent representation, computed by a Perceiver-based encoder that processes the reference mel spectrogram into a compact vector capturing the speaker's identity. A HiFi-GAN vocoder converts the final token sequence into a 24kHz audio waveform.

What Improved Over the Original XTTS

Version 2's changes are concentrated in exactly the areas that make cloning reliable in practice rather than just plausible in a demo: better speaker conditioning, support for multiple reference clips with interpolation between speakers, improved prosody and audio quality across the board, and meaningfully better inference stability compared to the original release. It also carries over the original's cross-language cloning capability — record a reference clip in one language and generate natural speech in a different one, while preserving the speaker's vocal identity — along with emotion and style transfer that happens automatically through the reference encoder rather than through separate, hand-tuned controls.

Independent testing has reported speaker similarity scores (SECS) around 0.59 in pure zero-shot conditions, improving to roughly 0.72 with about ten minutes of additional fine-tuning data — a useful benchmark if you're deciding whether zero-shot cloning alone will be good enough for your project or whether it's worth the extra step of fine-tuning on a specific voice.

The Coqui Shutdown, and What It Actually Means for You

This is worth understanding clearly before you build anything on XTTS-v2. Coqui AI — the Berlin-based startup founded by former Mozilla DeepSpeech engineers — shut down in January 2024, discontinuing its hosted Coqui Studio and Coqui API products entirely. The company folded without being acquired, and its original GitHub repository is no longer officially developed by the founding team.

In practice, this changes less than you'd expect. The code and model weights remain fully public and downloadable, and an active community fork (idiap/coqui-ai-TTS) has kept the project compatible with current Python and PyTorch versions, patching issues the original maintainers never got to. If you install the original TTS package and hit failures on a modern Python version, installing the community fork's package instead resolves most of them. The practical downside is that there's no official support channel anymore, and the one thing that did meaningfully change is licensing: Coqui previously sold a commercial license for XTTS-v2's non-commercial model weights, and since the company no longer exists, that paid commercial license option is effectively gone.

Getting Started with XTTS-v2

  1. Install the community-maintained package if you're on a recent Python version. pip install coqui-tts (the actively maintained fork) is generally more reliable than the original pip install TTS package on Python 3.10–3.12 with recent PyTorch.
  2. Load the model through the same API either way. from TTS.api import TTS followed by TTS("tts_models/multilingual/multi-dataset/xtts_v2", gpu=True) works identically regardless of which package you installed.
  3. Watch for the PyTorch 2.6 security restriction. Newer PyTorch versions block XTTS-v2's custom XttsConfig class under default weights_only=True settings — you'll need to explicitly allow it via torch.serialization.add_safe_globals before loading the model, or generation will fail at the loading step.
  4. Generate a cloned clip. tts.tts_to_file(text="your text", speaker_wav="reference.wav", language="en", file_path="output.wav") covers the basic zero-shot cloning workflow in a few lines.
  5. Check your VRAM before deploying. Basic inference needs a minimum of roughly 3GB of GPU VRAM; for stable, faster-than-real-time generation in production, 8GB or more avoids memory issues during longer synthesis runs.
  6. Fine-tune if zero-shot quality isn't quite enough. Community forks maintain the training scripts needed for fine-tuning, since the original Coqui team is no longer maintaining them officially — expect a more hands-on setup process than a fully supported commercial product would offer.

Tips for Better Results

  • Use a clean, well-recorded 6-second (or longer) reference clip. Cloning quality is directly sensitive to recording conditions — background noise or distortion in the reference audio degrades fidelity noticeably more than it would with some newer architectures.
  • Try multiple reference clips for a more stable identity. XTTS-v2's support for multiple speaker references and interpolation between them can produce a more consistent voice than relying on a single short clip, particularly for voices that are harder to capture cleanly in one recording.
  • Set expectations on cross-lingual quality by voice. Speaker similarity for voices that are less represented in the training data can score lower than for common voice profiles — test your specific target voice and language pair before committing to a production pipeline.
  • Confirm your commercial licensing path before shipping a paid product. Since Coqui's original commercial license is no longer being sold, review the current Coqui Public Model License terms carefully, or consider a model with a more permissive license (Apache 2.0 or MIT) if commercial deployment is the goal.
  • Prefer the community fork for anything beyond casual experimentation. Given the original repository's reduced maintenance, the actively patched fork is the more dependable base for a project with any longevity.

XTTS-v2 vs. Other Zero-Shot Cloning Models

XTTS-v2F5-TTSChatterbox Multilingual
ArchitectureGPT-2-style autoregressive + VQ-VAEDiffusion Transformer + flow matchingDiffusion-based
Reference clip length~6 seconds~5–15 seconds~5–10 seconds
Languages17Primarily English-focused23
Multiple reference clips / interpolationYesNot a core featureNot a core feature
Active official maintenanceNo (community fork only)YesYes
Weights licenseCoqui Public Model License (non-commercial)CC-BY-NCMIT

XTTS-v2's staying power comes from being an early, genuinely solid answer to multilingual zero-shot cloning at a time when few open alternatives existed — and its architecture and language coverage have held up well enough that it remains a standard comparison point even years after its own creator stopped existing.

Frequently Asked Questions

Is Coqui AI still developing XTTS-v2?

No. Coqui AI shut down as a company in January 2024. The code and model weights remain public, and an actively maintained community fork keeps the project working on current software, but there's no official Coqui support or development happening anymore.

Can I still use XTTS-v2 commercially?

It requires care. The underlying TTS library code is under the commercial-friendly Mozilla Public License 2.0, but the XTTS-v2 model weights themselves are released under the non-commercial Coqui Public Model License. Coqui previously sold a separate commercial license for the weights, but since the company no longer exists, that paid option is no longer available — review current license terms directly before any commercial use.

How much reference audio does XTTS-v2 need for cloning?

As little as 6 seconds, though independent testing shows speaker similarity improves further with additional fine-tuning data, such as around ten minutes of audio for a specific voice.

Why does installation sometimes fail on newer Python or PyTorch versions?

The original TTS package can run into compatibility issues on newer environments; installing the community-maintained coqui-tts fork resolves most of these, and PyTorch 2.6+ requires an explicit torch.serialization.add_safe_globals call to load XTTS-v2's custom configuration class.

How does XTTS-v2 compare to newer models like F5-TTS or CosyVoice3?

XTTS-v2 still offers competitive multilingual coverage and reliable cloning quality, but it lacks official ongoing development, and several newer models offer lower latency or more permissive licensing. It remains a strong choice for broad language support and community-tested stability rather than cutting-edge speed.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →