EchoraEchora
Back to Models

VALL-E X: One Sentence In, Any Language Out

September 1, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

VALL-E proved that a 3-second clip was enough to clone a voice within a single language. The obvious next question was harder: could that same short clip let someone speak a language they've never actually spoken? VALL-E X, Microsoft's cross-lingual follow-up, answers yes — take one sentence from a monolingual speaker, and it generates natural, fluent speech from that same voice in an entirely different language, without ever needing recordings of that person actually speaking it.

What Is VALL-E X?

VALL-E X is a cross-lingual neural codec language model, introduced by Microsoft researchers as a direct extension of the original VALL-E architecture. It inherits VALL-E's core in-context learning capability — the ability to pick up a voice from a short prompt without additional training — and extends it specifically to cross-lingual generation, trained on large-scale multilingual, multi-speaker speech data spanning multiple domains.

The practical result: provide one sentence in a speaker's native language, and VALL-E X generates speech from that same voice in a target language, while preserving the speaker's vocal identity, emotional tone, and even the acoustic background present in the original clip. No paired recordings of that speaker in both languages are required — the model never needs to have heard this specific person speak the target language at all.

How the Cross-Lingual Transfer Actually Works

The mechanism is a natural extension of VALL-E's phoneme-and-acoustic-token approach, adapted for the cross-lingual case. VALL-E X takes phoneme sequences derived from both the source and target text, along with the source speaker's acoustic tokens (extracted from their reference clip via a neural audio codec), and uses all of this as a combined prompt to generate acoustic tokens in the target language. Those output tokens are then decoded back into an actual speech waveform in the same way VALL-E's original architecture does.

Training doesn't require matched cross-lingual data from the same speakers — the model instead learns from a multilingual dataset where phoneme and acoustic token sequences from different languages are concatenated together, alongside a language ID signal that tells the model which language it's generating for at any given point. That language ID mechanism is also what gives VALL-E X a second, more playful capability: deliberately controlling accent, including generating Chinese speech with a distinctly English accent or vice versa, rather than only aiming for a fully native-sounding result.

Preserving More Than Just the Voice

As with the original VALL-E, working from real acoustic tokens rather than an abstracted speaker embedding means VALL-E X carries over more than vocal identity across the language switch — emotional tone and acoustic environment from the source clip tend to persist in the target-language output as well. This is a meaningful practical detail for use cases like dubbing or content localization, where losing a speaker's emotional delivery when switching languages would undermine the whole point of using their actual voice in the first place.

How VALL-E X Was Evaluated

Microsoft's evaluation used LibriSpeech alongside the EMIME and AISHELL-3 datasets to test English-Chinese cross-lingual generation in both directions — generating English speech prompted by Chinese speakers, and Chinese speech prompted by English speakers. The reported results demonstrated high-quality zero-shot cross-lingual synthesis, meaningfully reducing the foreign-accent artifacts that had been a persistent weak point in earlier cross-lingual TTS approaches, while the same underlying architecture was also shown capable of extending toward speech-to-speech translation tasks.

Getting Started with VALL-E X Today

As with the original VALL-E, Microsoft published the research but never released official code or pretrained model weights. Unlike that first model, though, the community response here produced something more immediately usable:

  1. Use the Plachtaa/VALL-E-X community reproduction. This open-source implementation trained its own model from scratch and — notably — actually released the resulting pretrained weights publicly, making it the more practical entry point compared to VALL-E reproductions that require training from zero.
  2. Expect coverage of three languages. The community model supports English, Chinese, and Japanese, with natural, expressive synthesis across all three and cross-lingual generation between them.
  3. Clone a voice from 3–10 seconds of reference audio. This is the practical range the reproduction was built around for zero-shot voice cloning before attempting cross-lingual generation.
  4. Know that the decoder was upgraded from the original design. The community implementation replaced the EnCodec decoder with a Vocos decoder specifically to improve audio quality, along with adding batched AR decoding for more stable generation results — genuine improvements over a literal reproduction of the original paper.
  5. Don't expect built-in speech-to-speech translation. The original Microsoft research described an S2ST extension, but the community reproduction doesn't include that module, since it requires separate training; pairing the model with a standard translation API (like Google Translate) to get source text into the target language before generation is the documented workaround.
  6. Experiment with accent control directly. Beyond straightforward cross-lingual cloning, try deliberately mismatching the language ID signal from the target output language to hear the accent-control capability in action — speaking one language with another's accent — a genuinely distinctive feature relative to most cloning models.

Tips for Better Results

  • Provide a clean single-sentence reference clip. Since the entire cross-lingual transfer hinges on that one source-language prompt, recording quality there has an outsized effect on the fidelity of the target-language output.
  • Test both directions of your specific language pair. Cross-lingual quality isn't necessarily symmetric — generating Language B from a Language A speaker and the reverse can perform differently, so validate the specific direction your project actually needs.
  • Use language ID deliberately if accent is part of your creative goal. This isn't just a side effect to tolerate — for content where a stylized accent is actually desirable (say, a character with a specific vocal quirk), the language ID mechanism is a genuine creative tool rather than only an accuracy setting.
  • Plan your translation step separately if you need full speech-to-speech translation. Since S2ST isn't bundled into the community reproduction, budget for an external translation step in your pipeline rather than expecting the model to handle source-to-target text conversion on its own.
  • Check licensing on the specific community reproduction you use. As with any unofficial reproduction of a research model, review the training data provenance and license terms of whichever implementation you adopt before commercial use.

VALL-E X vs. VALL-E vs. Other Multilingual Cloning Models

VALL-E XVALL-E (original)XTTS-v2
Core capabilityCross-lingual voice transferSame-language zero-shot cloningMultilingual cloning, 17 languages
Cross-lingual, same speakerYes, core design goalNot supportedYes
Accent controlYes, via language IDNot applicableNot a dedicated feature
Reference clip length~3–10 seconds~3 seconds~6 seconds
Official public weightsNoNoYes
Usable community reproductionYes, with released pretrained weightsRequires self-trainingN/A, officially released
Emotion/environment carryoverYes, across languagesYes, same languageAutomatic via reference encoder

VALL-E X's specific contribution is proving that cross-lingual voice transfer didn't need matched training data from the same speaker in both languages — a genuinely harder problem than same-language cloning, and one it solved using the same in-context learning approach that made the original VALL-E notable.

Frequently Asked Questions

Can I download official VALL-E X weights from Microsoft?

No. As with the original VALL-E, Microsoft published the research describing VALL-E X but never released official code or pretrained models. The community reproduction Plachtaa/VALL-E-X is the practical path to using this architecture today, and it does provide ready-to-use pretrained weights.

What languages does VALL-E X support?

Microsoft's original research focused evaluation on English and Chinese cross-lingual generation. The popular community reproduction extends coverage to English, Chinese, and Japanese.

Does VALL-E X need recordings of the same speaker in both languages?

No — that's the core problem it solves. It generates target-language speech in a speaker's voice using only a source-language reference clip, with no matched cross-lingual training data required for that specific speaker.

Can VALL-E X control accent deliberately?

Yes, through its language ID mechanism, which can be used both to minimize foreign accent for natural-sounding native speech, or deliberately applied to produce a specific accent effect, like speaking Chinese with an English accent.

Does VALL-E X include speech-to-speech translation?

Microsoft's original research described an extension toward this capability, but the popular community reproduction doesn't include it out of the box — pairing the model with a separate translation step is the documented workaround.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →