VALL-E: The Model That Started Treating Speech Like Text
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
GPT models don't understand language by regressing continuous signals — they predict the next token in a sequence, over and over, and that simple idea scales remarkably well. VALL-E, developed by Microsoft, asked what would happen if text-to-speech worked the same way: convert audio into discrete tokens, then let a language model predict them the way it would predict words. That reframing is why VALL-E is described as a neural codec language model rather than a conventional TTS system, and it's the architectural starting point that an entire generation of later voice-cloning models — VALL-E 2, VALL-E X, and beyond — built on directly.
What Is VALL-E?
VALL-E was introduced by Microsoft in a January 2023 paper, and its defining move was treating text-to-speech as a conditional language modeling task over discrete audio codes, rather than the continuous signal regression earlier neural TTS systems relied on. To make that possible, VALL-E uses EnCodec — a neural audio codec built on residual vector quantization — to convert real speech waveforms into discrete tokens that carry both speaker identity and acoustic characteristics, and which can later be decoded back into high-quality audio.
The model was trained on roughly 60,000 hours of English speech — hundreds of times more training data than TTS systems typically used at the time — and the payoff of that scale was a genuinely new capability: in-context learning for voice cloning. Give VALL-E just a 3-second acoustic prompt from a speaker it's never heard before, and it can generate new speech in that voice, the same way a large language model can pick up a pattern from a short prompt without any additional training.
How Speech Becomes a Language Modeling Problem
VALL-E's architecture is a hybrid of autoregressive (AR) and non-autoregressive (NAR) modeling, split across the layered structure of its codec tokens. EnCodec doesn't produce just one stream of tokens per audio frame — it produces eight codebook layers per frame, each capturing progressively finer acoustic detail. VALL-E's AR component predicts the first, most important codebook layer sequentially, conditioned on the phoneme sequence from the input text — this is the step that behaves most like a traditional language model generating one token at a time. The NAR component then fills in the remaining seven layers in parallel, conditioned on what the AR model already generated, trading some of the AR model's flexibility for meaningfully faster generation on the layers that matter less for overall intelligibility.
A Notable Side Effect: Emotion and Environment Carry Over
Because the model works from a genuine acoustic prompt rather than an abstracted speaker embedding, VALL-E was found to preserve more than just vocal identity from its 3-second reference clip — it also tends to carry over the speaker's emotional tone and even the acoustic environment (room reverb, background characteristics) present in the prompt itself. This wasn't an explicitly engineered feature so much as an emergent property of conditioning directly on real acoustic tokens rather than a compressed, identity-only representation — the same kind of emergent behavior that made in-context learning notable in text-based language models to begin with.
How VALL-E Performed
In its original evaluation against YourTTS — the strongest publicly available zero-shot TTS system at the time — VALL-E showed significantly better results on both speech naturalness and speaker similarity, using the LibriSpeech and VCTK benchmarks. It also demonstrated the ability to generate diverse outputs for the same input through different sampling-based decoding runs, rather than always collapsing to one deterministic result.
Getting Started with VALL-E Today
This is important to know upfront: Microsoft never publicly released VALL-E's training code or model weights, citing the potential for misuse given how convincing 3-second voice cloning can be — what exists publicly is the research paper and demo samples, not an official downloadable checkpoint. Practical access today comes through community efforts instead:
- Use an independent PyTorch reimplementation. Projects like
lifeiteng/vall-eon GitHub reproduce the architecture from the paper, but generally require you to train your own model on a dataset like LibriTTS rather than providing ready-to-use pretrained weights. - Consider VALL-E X community reproductions for ready-to-use weights. VALL-E X — a cross-lingual extension of the same core idea, also originally a Microsoft research project without a public official release — has a popular community reproduction (
Plachtaa/VALL-E-X) that does ship pretrained weights, making it a more immediately usable starting point if you want to hear the architecture family in action without training from scratch. - Set expectations for training if you go the from-scratch route. Given the original model's 60,000-hour training scale, reproducing comparable quality on your own hardware requires meaningful compute and a large, clean speech dataset — this is a research-grade undertaking, not a weekend project.
- Look to VALL-E 2 for the more usable descendant. Microsoft's follow-up work addressed VALL-E's known robustness and speed issues directly, and represents the more refined version of this same core approach — worth exploring specifically if production usability matters more to you than studying the original architecture.
Tips for Working with VALL-E
- Treat it as the architectural reference point, not a deployment-ready model. Given the lack of an official public release, most practical value today comes from understanding the neural-codec-language-model approach VALL-E pioneered, which shows up in various forms across many newer open-source cloning models.
- Expect emotional and environmental carryover from your prompt clip. If you're working with a reproduction and get output that unexpectedly carries the reference clip's background noise or emotional tone, that's a documented characteristic of this architecture family, not a bug in whichever implementation you're using.
- Budget for real training time if using a from-scratch reimplementation. Community implementations without pretrained weights require substantial data and compute to reach anything close to the original paper's reported quality.
- Check the reproduction's specific license and dataset provenance. Since these are community-trained models rather than an official release, licensing terms and training data sourcing vary by project — review the specific repository you're using before any commercial application.
VALL-E vs. Contemporary and Later Models
| VALL-E | YourTTS (contemporary baseline) | Later flow-matching models (E2 TTS, F5-TTS) | |
|---|---|---|---|
| Core approach | Discrete codec tokens + language modeling | Traditional neural TTS | Flow matching, non-autoregressive |
| Training data scale | ~60,000 hours | Considerably smaller | Comparable or smaller, different focus |
| Voice cloning prompt length | ~3 seconds | Longer typically required | ~5–15 seconds |
| Naturalness / similarity (reported) | Outperformed YourTTS | Baseline of comparison | Generally faster convergence, more stable alignment |
| Official public weights | No | Yes | Yes (F5-TTS); no (E2 TTS) |
| Emotion/environment carryover | Notable, emergent | Not a highlighted feature | Less emphasized |
VALL-E's real legacy isn't a checkpoint you can download — it's the reframing of TTS as a token-prediction problem that a large share of today's open-source cloning models, in one form or another, still build on.
Frequently Asked Questions
Can I download official VALL-E model weights from Microsoft?
No. Microsoft released the research paper and demo audio samples but never publicly released the training code or model weights, citing the risk of misuse from highly convincing short-clip voice cloning.
What does "neural codec language model" actually mean?
It means the model treats speech generation as predicting discrete audio tokens (derived from a neural audio codec like EnCodec) the same way a language model predicts text tokens, rather than directly regressing a continuous audio signal the way earlier TTS systems did.
How much reference audio does VALL-E need for voice cloning?
As little as 3 seconds, using in-context learning to generate new speech in that voice without any additional fine-tuning.
What's the difference between VALL-E and VALL-E X?
VALL-E X is a cross-lingual extension of the same core architecture, also originally a Microsoft research project without an official public release, but with a popular community reproduction that does provide pretrained, ready-to-use weights.
Is there a way to actually run VALL-E without training it myself?
Community reproductions like Plachtaa/VALL-E-X provide pretrained weights for the cross-lingual variant; reproductions of the original VALL-E, like lifeiteng/vall-e, generally require training your own model rather than downloading finished weights.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.