EchoraEchora
Back to Models

Descript Overdub: Voice Cloning That Lives Inside Your Edit, Not a Separate App

September 16, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Other voice cloning tools ask you to leave your editing software, generate a clip elsewhere, and drag it back in. Descript Overdub skips that trip entirely: it's a voice cloning engine built directly into Descript's text-based video and podcast editor, where you fix a flubbed line the same way you'd fix a typo — by editing the transcript. It's also one of the most tightly scoped cloning tools in the category, by design: Overdub only clones your own voice, and understanding why tells you a lot about what it's actually built for.

What Is Descript Overdub?

Descript is a text-based audio and video editor founded in 2017 by Andrew Mason, the former CEO of Groupon, built around a simple premise: editing a transcript should edit the underlying recording automatically, so cutting a sentence from the text cuts it from the audio or video too. In 2019, Descript acquired Lyrebird AI, a Montreal research team that had pioneered consumer voice cloning two years earlier, and folded its founders and technology into Descript as an in-house research group. Overdub is the commercial product built on that acquisition — voice cloning offered not as a standalone app, but as one feature inside an editor already used by outlets like NPR, Vice, and The New York Times for podcast and video production.

Functional Advantages of Descript Overdub

  • Fixes are typed, not re-recorded. Misread a line or need to swap out a word? Type the correction into the transcript and Overdub generates it in your cloned voice — no re-booking studio time or re-setting up a mic for one sentence.
  • Boundary-aware insertion into real recordings. When a generated word or phrase lands in the middle of an existing take, Overdub matches the tonal characteristics on both sides of the edit, so the seam isn't audible the way a hard cut would be.
  • Voice creation from audio you already have. Training no longer requires reading a fixed script for 10–30 minutes; Descript can build a voice from existing recordings — ideally from the same podcast or project you plan to use Overdub on.
  • A stock voice library as a fallback. If you don't want to clone your own voice at all, Overdub includes a library of licensed synthetic voices you can drop into a project the same way.
  • One workflow instead of two tools. Because cloning lives inside the same editor as your transcript, timeline, and filler-word removal, there's no export-import round trip between a voice tool and your editing software.

How Overdub's Cloning Architecture Actually Works

Overdub's engine is a direct descendant of Lyrebird AI's research, and it's built in two distinct stages rather than training a fresh model per user from scratch. The first stage trains a general speech model on recordings from thousands of different speakers, teaching it the broad patterns of human speech — phonetics, rhythm, prosody — independent of any one voice. The second stage adapts that general model to your specific voice using only your own sample audio, which is why a workable clone doesn't require the enormous, per-person datasets earlier voice synthesis approaches needed. The underlying generation itself relies on generative adversarial networks (GANs), pitting a generator against a discriminator during training to push the output toward higher-resolution, more natural-sounding audio.

In practice, Descript recommends at least 30 minutes of clean training audio, with quality continuing to improve as you approach 90 minutes — 10 minutes is the technical floor, but Descript's own guidance calls that a minimum, not a target. Processing typically takes 24–48 hours before a new voice becomes available. Once trained, output quality isn't uniform across every use case: Overdub is strongest correcting single words and short phrases spliced into real recordings, which is its core design goal, but pushing it to generate something longer — a fresh 30-second intro with no real audio to match against — tends to flatten tone and introduce a more robotic rhythm, a limitation Descript hasn't fully engineered around.

One more architectural choice is worth understanding on its own: on Free and Creator accounts, an Overdub voice ships with a vocabulary capped at roughly 1,000 common words, and typing outside that list produces gibberish rather than the word you intended — full, unrestricted vocabulary is reserved for higher-tier accounts. That's a product gate rather than a limit of the underlying model, but it means the trial experience and the production experience aren't actually the same technology tier. Overdub also currently works in English only, a scope constraint that hasn't changed since the feature's public release.

Tips for Better Descript Overdub Results

  • Record training audio in a quiet, acoustically dead room with an external mic. Descript's own guidance is explicit that background noise and mic quality have an outsized effect on how production-ready the resulting voice sounds.
  • Aim past the 10-minute floor. Since quality keeps improving up to roughly 90 minutes of training audio, treat 10 minutes as the bare minimum for a test, not the target for anything you'll publish.
  • Use Overdub for corrections, not full narration. It's built and tuned for fixing single words or short phrases inside a real recording — asking it to generate a long, standalone passage is where the output is most likely to sound synthetic.
  • Train from existing project audio when you can. Building a voice from recordings you already have skips the older script-reading step and can better match the tone of the specific content you're editing.
  • Confirm your plan's vocabulary limit before you rely on it for client work. A 1,000-word cap turning an unexpected brand name or technical term into gibberish is a bad surprise to discover mid-deadline rather than before you start.

Descript Overdub vs. ElevenLabs

Descript OverdubElevenLabs
Where it livesBuilt into a text-based video/audio editorStandalone platform and API
Whose voice you can cloneOnly your own — no exceptionsOwn voice or others', with documented consent
Core use caseFixing flubbed lines inside real recordingsFull narration, dialogue, and multi-voice production
Training data10–90 minutes, from a script or existing audio1–3 minutes (Instant) or 30 min–3 hrs (Professional)
Vocabulary limitsCapped on lower tiers, unlimited on higher tiersNot vocabulary-limited

The two tools aren't really competing for the same job. Overdub is a correction feature wrapped around an editor you're probably already using for something else — its value is in never leaving your project to fix one word. ElevenLabs is a dedicated voice generation platform built for producing new, standalone audio at scale. If your actual need is "clone a character voice for a story" or "narrate a full audiobook," Overdub's own-voice-only constraint rules it out entirely; if your need is "stop re-recording every time I flub a line in my podcast," that's precisely the problem it was built to solve.

Frequently Asked Questions

What is Descript Overdub?

Overdub is Descript's built-in voice cloning feature, based on technology from Lyrebird AI, that lets you generate new audio in your own cloned voice by typing text directly into a project's transcript.

Is Descript Overdub AI actually voice cloning, or just text-to-speech?

It's genuine voice cloning — Overdub trains a model on your specific voice from a sample recording, rather than using a fixed generic voice, and that trained model is what generates new speech from typed text.

Can Descript Overdub clone someone else's voice?

No. Overdub is restricted to cloning only the account holder's own voice, a deliberate consent safeguard rather than a technical limitation of the underlying model.

How much training audio does Descript Overdub voice cloning need?

Ten minutes is the technical minimum, but Descript recommends at least 30 minutes, with results continuing to improve up to around 90 minutes of clean source audio.

Why does my Descript Overdub voice only say some words correctly?

On lower-tier accounts, Overdub voices are capped at roughly a 1,000-word vocabulary — text outside that list won't generate properly. Unlimited vocabulary is available on higher-tier plans.

Generate Replacement Lines with Echora

To create new lines in a familiar voice through a similar workflow, upload a reference recording to Echora’s Voice Clone and enter your revised text, up to 5,000 characters. Generate speech that resembles the reference voice, then preview and download a separate audio clip to bring into your editor. Use only your own voice or a voice you have explicit permission to clone.

Generate a Replacement Voice Clip →