VoiceCraft: Edited Speech That Listeners Preferred Over the Real Recording
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
In VoiceCraft's own side-by-side naturalness comparison, human listeners preferred its edited speech over the actual original recording 48% of the time — essentially a coin flip between a real clip and one that had a word swapped out and regenerated by a model. That result is the clearest evidence of what VoiceCraft was actually built to do: not just generate speech from scratch, but edit existing audio convincingly enough that the seam disappears.
What Is VoiceCraft?
VoiceCraft is a token infilling neural codec language model developed by researchers at the University of Texas at Austin and Rembrand, released as fully open source — data, code, and model weights all publicly available on GitHub and Hugging Face. It's built as a Transformer decoder, and at release it achieved state-of-the-art results on both speech editing and zero-shot text-to-speech, evaluated specifically on in-the-wild audio: audiobooks, YouTube videos, and Spotify podcasts, rather than clean studio recordings. Both cloning an unseen voice and editing an existing recording require only a few seconds of reference audio.
The Technical Trick: Causal Masking Plus Delayed Stacking
This is worth understanding, because it's what makes editing — not just generating from scratch — actually possible. Standard autoregressive models generate strictly left to right, which makes them poorly suited to inserting new content into the middle of an existing sequence without regenerating everything after the edit point. VoiceCraft solves this with a two-step token rearrangement procedure: causal masking, borrowed from techniques that succeeded in joint text-image modeling, and delayed stacking, which handles the multi-codebook structure of neural speech codecs efficiently. Together, these let VoiceCraft perform infilling generation — conditioned on context from both directions — while still using an autoregressive generation process under the hood, rather than requiring the fully non-autoregressive approach that infilling models typically rely on.
Two Capabilities in One Model
Because of that architecture, VoiceCraft handles two genuinely distinct tasks through the same underlying mechanism. Speech editing lets you insert, delete, substitute, or make multi-span changes to an existing recording — correcting a word, replacing a phrase, or adjusting a sentence without re-recording anything, with the model regenerating just the affected span rather than the entire clip. Zero-shot voice cloning uses that same infilling capability in reverse: given a few seconds of an unseen speaker's voice, generate entirely new speech in that voice. Most TTS models are built for one of these tasks or the other; VoiceCraft treats them as the same underlying problem solved by the same model.
The REALEDIT Benchmark
To evaluate speech editing specifically, the VoiceCraft team built REALEDIT, described as the first realistic benchmark of its kind: 310 real-world examples sourced from audiobooks, YouTube videos, and Spotify podcasts, ranging from 5 to 12 seconds, covering insertions, deletions, substitutions, and multi-span edits of between 1 and 16 words. That 48% preference result — edited speech chosen over the genuine original recording in a side-by-side comparison — came directly from testing against this benchmark, which is part of why it's a meaningfully credible number rather than a cherry-picked demo clip.
Getting Started with VoiceCraft
The most straightforward path to trying VoiceCraft is through the project's Docker-based setup, which handles environment configuration for you; the inference_tts.ipynb notebook is the documented starting point for testing zero-shot TTS generation once your environment is running. If you're planning to fine-tune or train on your own data, the project provides separate setup and training documentation, with AdamW recommended specifically for fine-tuning stability over training from scratch. Model weights have been publicly available on Hugging Face since March 2024, so downloading a pretrained checkpoint rather than training from zero is the practical starting point for most use cases.
Tips for Better Results
- Use editing for surgical fixes, not full rewrites. VoiceCraft's editing strength is in targeted insertions, deletions, and substitutions of a phrase or short span — for wholesale content changes, standard zero-shot generation from scratch is the more appropriate mode.
- Keep your reference clip clean for cloning. Since only a few seconds of audio are needed, a clear, well-recorded sample of that length will outperform a noisier one of similar duration.
- Test on your actual audio type before assuming benchmark performance transfers directly. REALEDIT was specifically built from in-the-wild sources like podcasts and YouTube videos rather than clean studio audio, which is a meaningfully realistic test bed — but your own content's specific noise profile and recording quality are still worth validating against directly.
- Use Docker unless you have a specific reason to manage dependencies manually. Given how many components (codec, phoneme processing, Transformer decoder) need to align correctly, the Docker path avoids a substantial share of common setup issues.
- Check for VoiceCraft-X if you need languages beyond English. A follow-up model extends the same core editing-and-cloning approach to 11 languages, using a large language model for cross-lingual text processing without requiring phonetic pronunciation lexicons — worth checking if your project's language needs have grown past the original English-focused release.
VoiceCraft vs. Generation-Only Cloning Models
| VoiceCraft | F5-TTS | E2 TTS | |
|---|---|---|---|
| Speech editing (insert/delete/substitute in existing audio) | Yes, core capability | No, generation only | No, generation only |
| Zero-shot voice cloning | Yes, few seconds of reference | Yes, ~5–15 seconds | Yes, similar |
| Core mechanism | Autoregressive with causal masking + delayed stacking for bidirectional context | Non-autoregressive flow matching | Non-autoregressive flow matching |
| Public release | Yes, data, code, and weights all open | Yes, code and weights | No, research paper only |
| Benchmark focus | In-the-wild audio (audiobooks, YouTube, podcasts) | Standard TTS benchmarks | Standard TTS benchmarks |
VoiceCraft's specific niche is being one of the few open models that treats editing an existing recording and generating a new one as the same underlying problem — most models on this site are built purely for generation and have no mechanism for touching audio that already exists.
Frequently Asked Questions
Can VoiceCraft edit a recording without regenerating the whole clip?
Yes — that's its core design goal. Using causal masking and delayed stacking, VoiceCraft regenerates just the targeted span (an inserted, deleted, or substituted word or phrase) rather than requiring the entire clip to be recreated.
How much reference audio does VoiceCraft need?
Just a few seconds, for both zero-shot voice cloning of an unseen speaker and for editing an existing recording.
What is the REALEDIT benchmark?
It's a realistic speech-editing evaluation dataset the VoiceCraft team built specifically for this purpose, containing 310 real-world examples from audiobooks, YouTube videos, and podcasts, covering insertions, deletions, substitutions, and multi-span edits.
Is VoiceCraft free and open source?
Yes. Data, code, and pretrained model weights are all publicly available on GitHub and Hugging Face, with no restrictions beyond the project's stated ethical use expectations, which explicitly prohibit unauthorized speech generation or manipulation of someone's voice without consent.
Is there a multilingual version of VoiceCraft?
Yes. VoiceCraft-X, a follow-up model, extends the same unified editing-and-cloning approach to 11 languages, using a large language model for cross-lingual text handling instead of requiring separate phonetic lexicons per language.
Create a Similar Voice from Reference Audio
For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.