NaturalSpeech 3: Speech Isn't One Thing, So Why Generate It Like It Is?
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most TTS models treat speech as a single, tangled signal to generate all at once — content, tone, voice, and texture bundled together and produced in one pass. NaturalSpeech 3, the latest release in Microsoft's NaturalSpeech research line, starts from a different premise: pull those threads apart first, generate each one on its own terms, and then recombine them. It's a divide-and-conquer approach to speech synthesis, and it's what NaturalSpeech 3 is specifically built around.
What Is NaturalSpeech 3?
NaturalSpeech 3 is a zero-shot text-to-speech system developed by Microsoft Research Asia and Microsoft Azure Speech, in collaboration with researchers from the University of Science and Technology of China, the Chinese University of Hong Kong, and Zhejiang University. It's the latest entry in the NaturalSpeech research line — following the original NaturalSpeech's single-speaker, human-level quality benchmark and NaturalSpeech 2's scaling to zero-shot, multi-speaker, and singing synthesis via latent diffusion — and it tackles a problem neither predecessor was specifically designed to solve: speech quality still fell short on similarity and controllability, in part because content, prosody, timbre, and acoustic texture were all being modeled as one entangled representation.
NaturalSpeech 3's core move is factorization: decompose speech into separate subspaces representing distinct attributes, model each one individually, and combine them at the end. Two components make that possible — a purpose-built neural codec for the disentanglement, and a diffusion model built to generate each factorized attribute in turn.
FACodec: A Neural Codec Built to Separate, Not Just Compress
Most neural audio codecs are designed to compress speech into tokens as efficiently as possible, with no particular concern for what those tokens actually represent semantically. FACodec, the codec NaturalSpeech 3 introduces, is designed around the opposite goal: using Factorized Vector Quantization (FVQ), it deliberately decomposes a speech waveform into distinct subspaces for content, prosody, timbre, and acoustic detail, rather than one undifferentiated code stream.
Achieving genuinely clean separation between these subspaces — rather than each one leaking information the others should be responsible for — required several specific techniques layered into training: an information bottleneck to strip out attributes that don't belong in a given subspace, dedicated supervised losses that explicitly teach each subspace its intended attribute, adversarial training, and gradient reversal to further discourage cross-contamination between subspaces. On the decoding side, the timbre subspace is reintroduced through conditional layer normalization as the final waveform is reconstructed, rather than being flatly concatenated with everything else.
The Factorized Diffusion Model
Once speech can be reliably split into these attribute subspaces, NaturalSpeech 3 generates them with a diffusion model built specifically for that factorized structure — producing duration, prosody, content, and acoustic detail in sequence, each conditioned on the attributes already generated before it and on the phoneme-level text input. Duration is modeled explicitly as its own step (rather than folded into prosody) specifically because the system is non-autoregressive and needs an explicit timing signal to work from.
The practical payoff of this design is real, independent controllability: because each attribute has its own subspace and its own generation step, you can condition different attributes on different prompts. One reference clip can supply the timbre you want while a separate signal shapes the prosody — a level of granular control that a single, entangled representation simply can't offer, since adjusting anything in an undifferentiated system risks dragging every other attribute along with it.
The Benchmark Numbers
NaturalSpeech 3 posted a new state-of-the-art speaker similarity result, with a Sim-O score of 0.67 and an SMOS (similarity mean opinion score) of 4.01, against the strongest baseline system's 0.64 Sim-O and 3.69 SMOS at the time — a meaningful jump attributed specifically to FACodec's cleaner disentanglement of timbre from everything else. More broadly, the system outperformed prior state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility, and reported achieving on-par quality with actual human recordings. The architecture was further validated at scale, with results improving when pushed to roughly 1 billion parameters and 200,000 hours of training data — a meaningful jump from NaturalSpeech 2's already substantial 44,000-hour training set.
How NaturalSpeech 3 Fits the Series
Each generation in this research line solved a progressively different problem. The original NaturalSpeech established a rigorous, statistically testable definition of human-level quality and hit it on a single-speaker benchmark. NaturalSpeech 2 scaled that ambition to multi-speaker, zero-shot, in-the-wild generation — and, notably, zero-shot singing synthesis — by replacing autoregressive discrete-token generation with continuous latent diffusion. NaturalSpeech 3 addresses a different weakness entirely: even with strong zero-shot cloning, controllability and clean attribute separation remained limited, which is exactly what factorized codecs and factorized diffusion were built to fix.
NaturalSpeech 3 vs. NaturalSpeech 2 vs. NaturalSpeech
| NaturalSpeech 3 | NaturalSpeech 2 | NaturalSpeech (original) | |
|---|---|---|---|
| Core approach | Factorized codec (FACodec) + factorized diffusion | Continuous latents + latent diffusion | VAE-based, end-to-end |
| Problem solved | Independent attribute control, similarity | Multi-speaker, zero-shot, singing | Single-speaker human-level quality |
| Speaker similarity (Sim-O) | 0.67 (new state-of-the-art at release) | Not directly comparable metric | Not applicable (single-speaker) |
| Independent attribute control | Yes, via separate subspaces and prompts | Limited to overall speaker/style prompt | Not applicable |
| Training scale | 1B parameters, 200,000 hours | 44,000 hours | Single-speaker benchmark dataset |
| Official public release | No (research paper only) | No (community reimplementation exists) | No (community reimplementation exists) |
NaturalSpeech 3's specific contribution is proving that treating speech as separable attributes — rather than one indivisible signal — improves both quality and controllability simultaneously, extending the series' throughline of defining a specific weakness precisely and then building an architecture engineered to close exactly that gap.
Frequently Asked Questions
What does "factorized" mean in NaturalSpeech 3?
It refers to decomposing speech into separate, disentangled subspaces — content, prosody, timbre, and acoustic detail — each modeled and generated somewhat independently, rather than treating speech as one entangled representation generated all at once.
What is FACodec?
FACodec is the neural speech codec NaturalSpeech 3 introduces, using Factorized Vector Quantization to split a speech waveform into distinct subspaces for content, prosody, timbre, and acoustic detail, using techniques like information bottlenecks and adversarial training to keep those subspaces cleanly separated.
How does NaturalSpeech 3 compare to NaturalSpeech 2?
NaturalSpeech 2 focused on scaling zero-shot cloning to multi-speaker, in-the-wild data and singing synthesis using continuous latent diffusion. NaturalSpeech 3 builds on that scale but specifically targets similarity and independent controllability, using its factorized codec and diffusion approach to reach a new state-of-the-art speaker similarity score.
Can I download or run NaturalSpeech 3?
Not officially. Microsoft has not released public code or pretrained weights, and unlike NaturalSpeech 2, no well-established community reimplementation of NaturalSpeech 3's specific architecture is widely available as of this writing.
Does NaturalSpeech 3 achieve human-level quality like the original NaturalSpeech?
Its published results report achieving on-par quality with human recordings, extending the series' original human-level-quality ambition to a much harder zero-shot, multi-attribute-controllable setting rather than the original's single-speaker benchmark.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.