VALL-E 2: Two Fixes, and Suddenly It's Indistinguishable From a Human Voice
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Getting a model from "impressively close to human speech" to "reportedly indistinguishable from it" doesn't always take a new architecture — sometimes it takes fixing the two specific things that were quietly holding the original back. VALL-E 2, Microsoft's follow-up to VALL-E, made exactly two targeted changes to the same core design, and the result was described as the first zero-shot text-to-speech system to reach human parity — a claim tested directly against real human recordings, not just against other models.
What Is VALL-E 2?
VALL-E 2 builds directly on VALL-E's neural codec language modeling approach — the same core idea of treating speech generation as token prediction over discrete audio codes, using the same hybrid autoregressive (AR) and non-autoregressive (NAR) Transformer architecture, with a text embedding layer, a code embedding layer, and a code prediction layer shared across both components. What changed isn't the foundation — it's two specific mechanisms layered on top, each targeting a concrete weakness the original model had.
The Two Improvements That Made the Difference
Repetition Aware Sampling. VALL-E's original decoding process used straightforward random sampling, which could occasionally produce unstable output or fall into an infinite loop, repeating the same sound or phrase indefinitely. VALL-E 2 fixes this by tracking token repetition in the decoding history and adaptively switching between random and nucleus sampling at each step — if the repetition ratio crosses a defined threshold, the model swaps in a randomly resampled token from the original probability distribution instead of continuing down the repetitive path. The result is meaningfully more stable decoding without sacrificing the natural variation that sampling-based generation provides.
Grouped Code Modeling. Rather than modeling every individual codec code as a separate step, VALL-E 2 groups codec codes together and models each group within a single frame during the AR generation process. This shortens the effective sequence length the model has to process, which both speeds up inference and directly addresses the difficulty long sequences created for the original architecture.
Neither change touches the fundamental neural-codec-language-model idea VALL-E introduced — they're targeted engineering fixes to decoding stability and sequence length, and together they were enough to move the needle from "very good" to a specific, testable human-parity claim.
The Human Parity Results, in Context
Evaluation used two distinct test sets, and the harder one is the more telling result. On LibriSpeech test-clean, using 40 distinct speakers with a reference utterance as the cloning prompt for each, VALL-E 2 surpassed the original VALL-E on both speaker similarity (SMOS) and speech quality (CMOS) scores. On VCTK — a considerably harder benchmark with 60 speakers covering much more diverse accents, making zero-shot cloning meaningfully more difficult — VALL-E 2 not only surpassed VALL-E again, but matched or exceeded the actual ground-truth human recordings, using only a 3-second prompt per speaker. That specific result, on the harder and more accent-diverse dataset, is the basis for the human-parity claim.
The paper also reported that VALL-E 2 could reliably synthesize complex sentences that had been a known challenge for the original model — sentences involving unusual phrasing or difficult phoneme sequences that previously triggered the instability Repetition Aware Sampling was specifically designed to fix.
Why You Can't Download VALL-E 2
This is the most important practical thing to understand about this model, and Microsoft has been unusually direct about it: VALL-E 2 was deliberately not released as a public model or product, with researchers citing exactly the misuse risks you'd expect from a system this convincing — voice spoofing and impersonation of a specific speaker without consent. The research paper itself frames its experiments around the assumption that users consent to being the target speaker in any synthesis, which is a notably explicit ethical guardrail for a research publication to state directly, and reflects how seriously the team treated the implications of what they'd built.
Practically, that means there's no official checkpoint, and — unlike VALL-E X, which has a popular, fully usable community reproduction with released pretrained weights — no comparably complete open-source recreation of VALL-E 2's specific Repetition Aware Sampling and Grouped Code Modeling techniques has emerged as of this writing. What's publicly available is the research paper, Microsoft's own project page with demo audio samples, and evaluation frameworks like ELLA-V that were used alongside LibriSpeech and VCTK to assess the model's handling of complex generation tasks.
What This Means If You're Evaluating VALL-E 2
If your interest is understanding the state of the art in neural codec language models, VALL-E 2's contribution is genuinely worth studying: it demonstrates that decoding-strategy fixes, not just bigger models or more data, can be the difference between "impressive" and "human parity" on a rigorous, accent-diverse benchmark. If your interest is actually deploying something, the practical path runs through models that did ship — VALL-E X's community reproduction for cross-lingual work, or entirely separate open-source cloning models (F5-TTS, Chatterbox, CosyVoice3) for production use, since VALL-E 2 itself was never made available to run.
VALL-E 2 vs. VALL-E vs. VALL-E X
| VALL-E 2 | VALL-E (original) | VALL-E X | |
|---|---|---|---|
| Core improvement | Repetition Aware Sampling + Grouped Code Modeling | Original neural codec language model concept | Cross-lingual extension of VALL-E |
| Decoding stability | Meaningfully improved, fixes infinite-loop issue | Known instability on some inputs | Inherits VALL-E's approach |
| Claimed result | Human parity on LibriSpeech and VCTK | Outperformed YourTTS baseline | High-quality cross-lingual transfer |
| Reference prompt length | ~3 seconds | ~3 seconds | ~3–10 seconds |
| Official public release | No | No | No |
| Usable community reproduction | None widely available | Requires self-training | Yes, with released pretrained weights |
VALL-E 2's real significance is as a research milestone rather than a tool you can pick up and use — it's the clearest evidence yet that the neural-codec-language-model approach VALL-E started can, with the right decoding refinements, produce zero-shot speech that listeners can't reliably distinguish from a real recording.
Frequently Asked Questions
Can I download or use VALL-E 2?
No. Microsoft has not released official code or pretrained weights, and unlike VALL-E X, no widely available community reproduction of VALL-E 2's specific techniques exists as of this writing. What's public is the research paper and demo audio samples.
What does "human parity" mean in this context?
It means that on specific benchmark evaluations — LibriSpeech and the more accent-diverse VCTK dataset — listener ratings of VALL-E 2's generated speech matched or exceeded ratings of the actual ground-truth human recordings, using only a 3-second cloning prompt per speaker.
What are Repetition Aware Sampling and Grouped Code Modeling?
Repetition Aware Sampling adaptively adjusts the decoding process based on token repetition history to prevent unstable or looping output. Grouped Code Modeling groups codec codes together to shorten effective sequence length, improving both inference speed and the model's handling of longer sequences.
Why did Microsoft decide not to release VALL-E 2?
Researchers cited the risk of misuse given how convincing the cloning results are, specifically mentioning concerns like voice spoofing and impersonating a specific speaker without their consent.
How is VALL-E 2 different from VALL-E X?
VALL-E 2 improves the original VALL-E's same-language decoding stability and speech quality through Repetition Aware Sampling and Grouped Code Modeling. VALL-E X, a separate line of work, extends VALL-E specifically for cross-lingual voice transfer, and unlike VALL-E 2, has a usable community reproduction with released pretrained weights.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.