EchoraEchora
Back to Models

NaturalSpeech 2: It Can Make a Voice Sing That Has Never Actually Sung

September 2, 2026
•
7 min read

Here's the detail that stands out most about NaturalSpeech 2: give it a few seconds of someone simply talking — no singing, no melody, nothing musical at all — and it can generate that same voice singing an entirely different piece, in tune and in time. Microsoft's follow-up to the original NaturalSpeech was built to solve a much harder problem than its predecessor: scaling human-level quality beyond a single studio-recorded voice, out to arbitrary speakers, styles, and — as it turned out — an entirely different mode of vocal performance altogether.

What Problem Was NaturalSpeech 2 Actually Solving?

The original NaturalSpeech proved human-level quality was achievable on a single-speaker, studio-quality benchmark. Scaling that same quality to large, diverse, in-the-wild datasets — covering different speaker identities, prosodies, and styles like singing — turned out to be a genuinely different problem. Most large-scale TTS systems at the time tackled this by quantizing speech into discrete tokens and generating them one by one with an autoregressive language model — the same broad approach VALL-E made famous. That approach, though, tends to suffer from unstable prosody and a tendency to skip or repeat words, along with voice quality that degrades as the diversity of training data grows.

NaturalSpeech 2 takes a structurally different path: instead of discrete tokens predicted autoregressively, it uses a neural audio codec with residual vector quantizers to produce continuous latent vectors, and generates those vectors using a latent diffusion model conditioned on the input text — the same non-autoregressive philosophy its predecessor used, applied to a much harder, more diverse data setting.

The Speech Prompting Mechanism Behind Zero-Shot Generation

To make zero-shot generation work reliably at this scale, NaturalSpeech 2 introduces a dedicated speech prompting mechanism, designed specifically to enable in-context learning inside both the latent diffusion model and a separate duration/pitch predictor. In practice, this means a short reference clip doesn't just condition the final audio generation step — it also informs how the model predicts timing and pitch contour for the new content, which is a meaningfully different design from simply appending a speaker embedding to a single stage of the pipeline.

Zero-Shot Singing Synthesis, Explained

This is the capability referenced in NaturalSpeech 2's own subtitle — "natural and zero-shot speech and singing synthesizers" — and it's worth understanding how it actually works rather than treating it as a black-box party trick.

The model was scaled to a combined training set of roughly 44,000 hours of speech and singing data, with the singing portion assembled from web-crawled vocal recordings — background music and accompaniment removed with a dedicated audio-separation process, and misaligned samples filtered out using an ASR-based check — totaling around 30 hours of singing data, upsampled and blended into the broader speech training set.

To actually generate a singing performance, the model borrows the ground-truth pitch and duration contour from an existing singing reference, then applies a chosen voice prompt to render that melody and timing in a specific singer's timbre. The genuinely notable finding: that voice prompt doesn't have to come from a singing recording at all — a spoken-voice prompt is enough for NaturalSpeech 2 to generate a novel singing voice in that speaker's timbre, extending zero-shot generation into a performance mode the reference speaker was never actually recorded doing.

How It Performed

Across zero-shot evaluation on unseen speakers, NaturalSpeech 2 outperformed previous large-scale TTS systems by a substantial margin on prosody and timbre similarity, robustness, and overall voice quality — with particularly strong prosody similarity to the reference prompt shown across LibriSpeech and VCTK test conditions, and results that held up meaningfully well even from prompts as short as 3 seconds.

Getting Started with NaturalSpeech 2 Today

As with the original NaturalSpeech, Microsoft did not release official code or pretrained weights for NaturalSpeech 2 — what's public is the research paper and audio samples on the project's demo page. Community efforts fill the gap, though with some important caveats:

  1. Look at the lucidrains/naturalspeech2-pytorch reimplementation. This is an actively developed PyTorch implementation of the described architecture, built around EnCodec as the neural codec and using denoising diffusion for the generative process.
  2. Understand this is a from-scratch training framework, not a pretrained model. Like most community reproductions of research-only releases, you'll need your own dataset and training run — there's no official or community-trained checkpoint offering the paper's reported quality out of the box.
  3. Check the project's own development status before relying on it. Its documentation lists ongoing work items — including refinements to duration/pitch prediction during training and additional attention-mechanism improvements — so treat it as an evolving research tool rather than a finished, stable release.
  4. Expect to assemble your own singing dataset if that's your specific goal. Since NaturalSpeech 2's singing capability came from carefully sourced and filtered web-crawled data mixed into a much larger speech corpus, replicating that specific result requires similarly deliberate data collection, not just training on a generic speech dataset.

Tips for Understanding NaturalSpeech 2

  • Separate the two headline claims when evaluating it. Zero-shot speech cloning at scale and zero-shot singing synthesis from a speech-only prompt are two distinct, individually notable results — don't conflate them when deciding what's actually relevant to your use case.
  • Recognize the architectural bet against discrete-token language modeling. NaturalSpeech 2's choice of continuous latent diffusion over the discrete-codec-plus-autoregressive-LM approach used by contemporaries like VALL-E is a direct response to the instability and word-skipping issues that approach can produce — a useful comparison point if you're evaluating architectures more broadly.
  • Take prompt length seriously if you're studying or replicating it. The research specifically measured how prosody similarity changes with prompt length (testing 3, 5, and 10-second prompts), and longer prompts generally produced better-matched results — a detail worth accounting for if adapting similar techniques.
  • Treat the responsible AI framing as a real constraint, not boilerplate. Given how effective this approach is at replicating vocal identity in an entirely different performance context (singing), the consent and detection safeguards Microsoft describes are directly relevant to any responsible use of similar techniques, not just legal language to skim past.

NaturalSpeech 2 vs. NaturalSpeech vs. VALL-E

NaturalSpeech 2NaturalSpeech (original)VALL-E
Core approachContinuous latents + latent diffusionVAE-based, end-to-endDiscrete codec tokens + autoregressive LM
Best-suited scenarioMulti-speaker, zero-shot, in-the-wild, singingSingle-speaker, studio-qualityZero-shot cloning from short prompts
Singing synthesisYes, zero-shot, even from a speech-only promptNot a focusNot a focus
Training scale~44,000 hours (speech + singing)Smaller, single-speaker benchmark~60,000 hours
Known robustness issuesDesigned specifically to avoid discrete-token instabilityNot autoregressive, so not applicableRepetition/instability addressed later in VALL-E 2
Official public releaseNo (research demo only)No (community reimplementation only)No

NaturalSpeech 2's real contribution is showing that the discrete-token, autoregressive path VALL-E popularized wasn't the only way to scale zero-shot TTS — a continuous latent diffusion approach could match or exceed it on robustness and quality, while additionally unlocking a genuinely novel capability in cross-modal vocal performance transfer.

Frequently Asked Questions

Can NaturalSpeech 2 really make someone sing who has never sung before?

Yes, according to the published research — generating a novel singing performance in a speaker's timbre using only a spoken-voice prompt is one of its specifically highlighted, tested capabilities.

How is NaturalSpeech 2 different from NaturalSpeech?

The original NaturalSpeech focused on and was evaluated for single-speaker, studio-quality synthesis. NaturalSpeech 2 targets the much harder problem of multi-speaker, zero-shot, in-the-wild synthesis at scale, using a latent diffusion approach instead of the original's VAE-based design.

Why did Microsoft use diffusion instead of an autoregressive language model like VALL-E?

Autoregressive generation over discrete speech tokens can suffer from unstable prosody and word skipping/repetition, especially at scale. NaturalSpeech 2's continuous latent diffusion approach was designed specifically to avoid those failure modes.

Does NaturalSpeech 2 raise voice-cloning consent concerns?

Yes, and Microsoft addresses this directly — the research was conducted under the assumption that users consent to being the target speaker, and real-world deployment is described as needing a consent protocol and synthesized-speech detection safeguards.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →