EchoraEchora
Back to Models

Voicebox: A Model Built to Edit Speech, Not Just Generate It From Scratch

September 7, 2026
•
5 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

An autoregressive TTS model can only extend a clip from wherever it stopped — it can't reach back and fix a word in the middle without regenerating everything after it. Voicebox, Meta's research model from FAIR, was built around a fundamentally different task: given surrounding audio context and a target text, fill in the missing piece, anywhere in the clip, the way you'd use inpainting to edit the middle of an image rather than only paint past its edge.

What Is Voicebox?

Voicebox is a non-autoregressive flow-matching model trained on more than 50,000 hours of unfiltered, unenhanced speech — drawn from public-domain audiobooks across English, French, German, Spanish, Polish, and Portuguese. Rather than training it directly on the many separate tasks TTS systems typically specialize in, Meta trained it on a single, more general objective: text-guided speech infilling. That one task turned out to subsume most of what people actually want from a speech model, through the same kind of in-context learning that made GPT useful for tasks it was never explicitly trained on.

Why Infilling Changes What's Possible

Because Voicebox is trained to reconstruct a masked span of audio using both the surrounding speech and the target text, it can modify any part of a given sample — not just the end of a clip the way autoregressive models are structurally limited to. That flexibility is what lets a single model, trained on one task, generalize into several practically distinct capabilities: zero-shot text-to-speech synthesis from as little as 2 seconds of reference audio, cross-lingual style transfer (reading a passage in a different one of its six supported languages while preserving the reference speaker's style), noise removal, content editing, and diverse resampling of a given utterance — all as instances of the same underlying infilling task rather than separate specialized modes bolted on afterward.

The Technique Behind It: Conditional Flow Matching

Voicebox's non-autoregressive core is built on Conditional Flow Matching (CFM) — a method that transforms a simple starting distribution (Gaussian noise) into the target speech distribution over a continuous time variable, conditioned on both frame-wise linguistic information and the masked audio context. This is meaningfully different from a diffusion model's typical formulation, and Meta's own comparisons at the time showed it outperforming VALL-E on zero-shot TTS intelligibility and audio similarity, while generating speech roughly 20 times faster than the leading autoregressive models of that period.

Flow matching as a technique didn't stay contained to Voicebox — the same core idea underpins the approach later used by E2 TTS and F5-TTS, both of which build on the same "pad or mask, then generate" philosophy Voicebox helped establish, and it's a close conceptual relative of the flow-matching decoders used in models like CosyVoice2 and CosyVoice3. If you've used any of those, you've used a direct descendant of the approach Voicebox demonstrated at scale.

Voicebox vs. Its Open-Source Architectural Descendants

VoiceboxE2 TTSF5-TTS
Core techniqueConditional flow matchingFlow matching, flat U-Net TransformerConditional flow matching + Diffusion Transformer
Training task framingGeneral text-guided speech infillingText-audio infilling (no duration model/aligner)Same core approach, refined for alignment robustness
Zero-shot prompt length~2 secondsSimilar~5–15 seconds
Task generalityTTS, denoising, editing, resampling from one modelPrimarily TTS-focusedPrimarily TTS-focused
Public releaseNo (research paper only)No (community reproduction via F5-TTS repo)Yes, on GitHub and Hugging Face
Languages6English-focusedPrimarily English-focused

Voicebox's lasting significance is more about demonstrating that a single, general infilling objective — trained at genuine scale — could match or beat specialized systems across several distinct speech tasks at once, a template a meaningful share of the field has continued building on since.

Frequently Asked Questions

What does "speech infilling" mean?

It means generating a missing or masked portion of an audio clip using the surrounding context and a target text, which lets the model edit any part of a sample rather than only extend from its endpoint the way autoregressive models are limited to.

How is Voicebox related to models like E2 TTS and F5-TTS?

Both are architectural descendants built on the same conditional flow-matching approach Voicebox demonstrated — E2 TTS shares its infilling philosophy directly, and F5-TTS refined that same lineage with additional improvements for training stability and alignment.

How many languages does Voicebox support?

Six, based on its training data: English, French, German, Spanish, Polish, and Portuguese.

Did Meta address misuse risk with Voicebox?

Yes. Alongside the generative model, Meta's research specifically describes building a classifier capable of distinguishing authentic speech from Voicebox-generated speech, as a documented safeguard against the model's misuse potential.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.

Create a Voice Clone →