EchoraEchora
Back to Models

VITS: The Architecture Quietly Running Under a Huge Share of Open-Source TTS

September 9, 2026
•
8 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Plenty of well-known open-source voice models turn out to share the same ancestor without most users ever realizing it — Piper's ONNX voices, MeloTTS's multilingual backbone, YourTTS's speaker-conditioned cloning, and the SoVITS half of GPT-SoVITS are all built directly on the architecture this page is about. VITS is the classic end-to-end TTS system that proved a single-stage, parallel-generation model could actually match the quality of the older two-stage pipelines — and that proof is why so much of the open-source TTS ecosystem still traces back to it.

What Is VITS?

VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) was introduced by Jaehyeon Kim, Jungil Kong, and Juhee Son at Kakao Enterprise, published at ICML 2021. It's an end-to-end speech synthesis model that predicts a waveform directly from an input text sequence in a single stage, released under the MIT license. Before VITS, end-to-end models capable of single-stage training and parallel sampling existed, but their output quality reliably fell short of two-stage systems — a separately trained acoustic model feeding a separately trained vocoder. VITS closed that gap directly, and its official evaluation showed it outperforming the strongest publicly available two-stage models on mean opinion score using the LJSpeech dataset.

Functional Advantages of VITS

  • Genuinely end-to-end. Text goes in, a waveform comes out, with no separately trained vocoder stage to manage — a real architectural simplification compared to the two-stage pipelines it was built to outperform.
  • Parallel, non-autoregressive generation. Because VITS doesn't generate audio token by token in sequence, inference is fast relative to autoregressive alternatives.
  • A stochastic duration predictor. The same input text can be rendered with different valid rhythms and pacing across generations, rather than always producing an identical, deterministic result.
  • Competitive quality against two-stage systems, validated through mean opinion score evaluation on both single-speaker (LJSpeech) and multi-speaker (VCTK) benchmarks.
  • Native multi-speaker support through speaker embeddings, without requiring a fundamentally different architecture for the multi-speaker case.
  • A permissive MIT license and a large enough footprint in the field that it became a foundation many later projects built directly on top of, rather than starting their own acoustic architecture from scratch.

Under the Hood: The Conditional VAE Architecture

VITS is structured as a conditional variational autoencoder made up of three cooperating pieces: a posterior encoder, a decoder, and a conditional prior. The conditional prior is where a Transformer-based text encoder and a stack of normalizing-flow coupling layers work together to model spectrogram-based acoustic features from the input text. The decoder then reconstructs the actual waveform through a stack of transposed convolutional layers, following a similar design philosophy to the HiFi-GAN vocoder rather than requiring HiFi-GAN as a genuinely separate downstream stage.

What makes this combination work as well as it does is the training approach: variational inference augmented with normalizing flows, combined with an adversarial training process borrowed from GAN-style methods, together improving the expressive power of the underlying generative model beyond what a standard VAE alone typically achieves.

The Stochastic Duration Predictor

This is VITS's most distinctive design choice, and it addresses a real limitation of most earlier TTS systems: the same sentence, spoken by a real person twice, doesn't come out with identical timing and rhythm every time. VITS models this directly with a stochastic duration predictor, using uncertainty modeling over latent variables to synthesize speech with genuinely diverse rhythms from the same input text — rather than forcing a single deterministic duration for every phoneme regardless of context. This is what gives VITS output a natural one-to-many quality: one piece of text, multiple valid, differently paced renditions, matching how spoken language actually behaves.

Getting Started with VITS

  1. Grab an official pretrained checkpoint from Kakao Enterprise on Hugging Face. kakao-enterprise/vits-ljs is the single-speaker model trained on LJ Speech; kakao-enterprise/vits-vctk is the multi-speaker model trained on the VCTK dataset (109 speakers, various English accents) — pick based on whether you need one consistent voice or multiple speakers.
  2. Use the Hugging Face Transformers integration for the simplest setup. VITS is supported directly via the VitsModel and VitsTokenizer classes in the transformers library — from transformers import VitsModel, VitsTokenizer followed by loading either checkpoint by name gets you to inference in a few lines, with no separate vocoder to configure.
  3. Use the original jaywalnut310/vits GitHub repository if you want to train or fine-tune. This is the official reference implementation from the paper's authors, and the starting point most derivative projects (Piper, YourTTS, and others) began from when adapting VITS for their own specific use case.
  4. Check for a language- or dataset-specific community checkpoint before training your own. Given how widely VITS has been retrained on other languages and datasets by the community, searching Hugging Face for an existing checkpoint in your target language is often faster than training from scratch.

The VITS Family Tree: What's Actually Built on It

Piper pairs a VITS-based model with ONNX export and espeak-ng phonemization for ultra-lightweight, CPU-only edge deployment. MeloTTS builds on VITS and VITS2 specifically, layering in BERT-based linguistic features for better prosody. YourTTS modifies the VITS foundation with a dedicated speaker encoder and Speaker Consistency Loss to add zero-shot multi-speaker and voice-conversion capability. GPT-SoVITS uses a VITS-based system (the "SoVITS" half of its name) for the acoustic and voice-conversion side of its hybrid pipeline, paired with a GPT-based component for semantic modeling. None of these are copies of VITS — each took the same core architecture and modified it for a specific, different goal, which is exactly the kind of adaptability that makes an architecture genuinely foundational.

Tips for Better Results

  • Use the LJSpeech checkpoint for a single consistent voice, VCTK for multi-speaker needs. Don't default to one checkpoint without checking which matches your actual project's speaker requirements.
  • Expect natural variation across generations, not a bug. Because of the stochastic duration predictor, running the same text through VITS multiple times will produce slightly different pacing each time — this is the architecture working as designed, not inconsistency to debug.
  • Look at a derivative project first if your need is more specific than base TTS. If you need voice cloning, extreme lightweight deployment, or Chinese-specific tuning, one of VITS's direct derivatives (YourTTS, Piper, GPT-SoVITS) has likely already solved that specific problem better than adapting the base architecture yourself would.
  • Check VITS2 if training efficiency and quality improvements matter to your project. Kong and colleagues published a direct follow-up specifically targeting quality and efficiency gains over the original through adversarial learning and architecture refinements.

VITS vs. VITS2 vs. a Flow-Matching Alternative

VITSVITS2F5-TTS
Core approachConditional VAE + normalizing flows + adversarial trainingSame foundation, refined architecture and trainingDiffusion Transformer + conditional flow matching
Generation modeNon-autoregressive, parallelNon-autoregressive, parallelNon-autoregressive
Duration modelingStochastic duration predictorImproved duration modelingNo explicit duration model
Released202120232024
LicenseMITCheck specific implementationCC-BY-NC (non-commercial)
Ecosystem of derivativesExtremely largeSmaller, more recentGrowing, newer paradigm

VITS's lasting significance isn't that it's still the state of the art on its own — several newer architectures now outperform it on raw benchmark numbers — but that its core design proved influential and adaptable enough to become the literal foundation under a meaningful share of the open-source TTS tools people actually use today.

Frequently Asked Questions

Is VITS still worth using directly, or should I use a derivative?

For general research or a straightforward single-speaker or multi-speaker TTS need, the base VITS checkpoints remain solid. For more specific requirements — voice cloning, extreme lightweight deployment, or language-specific tuning — one of its derivatives (YourTTS, Piper, MeloTTS, GPT-SoVITS) has usually already adapted the architecture for that exact purpose.

What's the difference between VITS and VITS2?

VITS2 is a direct follow-up from an overlapping set of authors, specifically targeting improved quality and training efficiency over the original through adversarial learning refinements and architecture changes, while keeping the same core conditional VAE philosophy.

Why does VITS output sound slightly different each time for the same text?

This is intentional, coming from the stochastic duration predictor, which models the natural one-to-many relationship between text and speech — the same sentence can be spoken with different valid rhythms and pacing, and VITS deliberately reflects that rather than forcing identical output every time.

Is VITS free for commercial use?

Yes, the original VITS release is under the MIT license, one of the more permissive options — though if you're using a specific derivative or community-trained checkpoint, check that project's own license separately, since it may differ from the base model's terms.

Does VITS support multiple speakers?

Yes, through speaker embeddings — the official kakao-enterprise/vits-vctk checkpoint demonstrates this directly, trained across 109 speakers with a range of English accents.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →