Real-Time-Voice-Cloning: The Open-Source Project That Taught a Generation How Voice Cloning Actually Works
- What Is Real-Time-Voice-Cloning?
- Functional Advantages of Real-Time-Voice-Cloning
- How SV2TTS's Three-Stage Architecture Actually Works
- Getting Started with Real-Time-Voice-Cloning Locally
- Tips for Better Results with Real-Time-Voice-Cloning
- Real-Time-Voice-Cloning vs. Newer Open-Source Voice Cloning
- Frequently Asked Questions
- Generate New Speech from a Reference Voice in Echora
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Long before commercial platforms turned voice cloning into a monthly subscription, one free project taught a generation of developers how it actually worked under the hood. Real-Time-Voice-Cloning, built by Corentin Jemine and released on GitHub, clones a voice from a short sample using a three-stage pipeline called SV2TTS. It's been forked, studied, and built on so many times that it's effectively the reference point for neural voice cloning. It's not what you'd reach for if you want production-grade output today, but running it and understanding how it's built is close to a prerequisite for understanding everything that came after it.
What Is Real-Time-Voice-Cloning?
Real-Time-Voice-Cloning is a free, open-source voice cloning project created by Corentin Jemine (GitHub handle CorentinJ), built around SV2TTS — a three-stage technique for cloning a voice from a short audio sample, originally developed by a Google research team in 2018. Its tagline — "clone a voice in 5 seconds to generate arbitrary speech in real-time" — undersold how influential it would become: it has accumulated over 60,000 GitHub stars and 9,400 forks, making it one of the most widely used and referenced voice cloning codebases ever published. Its legacy extends past the repo itself — the encoder component was spun off into an independent package called Resemblyzer, now maintained under the Resemble AI organization (the company Jemine went on to work for), and the underlying technique was adapted into MockingBird, a widely used Mandarin-language fork built specifically for Chinese voice cloning.
Functional Advantages of Real-Time-Voice-Cloning
- Completely free and runs entirely on your own hardware. No account, no API key, no per-generation cost — everything runs locally once you've installed the dependencies and downloaded the pretrained models.
- A genuinely educational architecture. Because the encoder, synthesizer, and vocoder are separate, inspectable Python modules rather than one opaque model, it's one of the more approachable codebases for understanding how neural voice cloning actually works internally.
- The transfer-learning trick that made short-sample cloning possible in the first place. This project is a working implementation of the technique that let voice cloning generalize to a speaker the model had never heard during training — the conceptual foundation most later commercial cloning tools build on.
- A real toolbox, not just training scripts.
demo_toolbox.pyprovides a GUI for recording or loading a voice sample, generating speech, and inspecting the resulting embeddings, whiledemo_cli.pyoffers a scriptable command-line path for the same pipeline. - An ecosystem of forks and derivatives to learn from. Because it's popular and permissively licensed, community forks like MockingBird extend it to new languages, giving you working reference implementations beyond the original English-focused version.
- Pretrained models available without needing your own training run. You can clone a voice using the author's published pretrained weights immediately, reserving full retraining for cases where you actually need a different base dataset or language.
How SV2TTS's Three-Stage Architecture Actually Works
The core idea behind SV2TTS, and the reason this project could clone a voice from just a few seconds of audio back in 2019 when most TTS systems needed hours of transcribed speech per speaker, is a specific transfer-learning trick: separate the problem of "what does this voice sound like" from the problem of "how do you make a TTS model read text in that voice," and train each half on the dataset that's actually easy to get.
The encoder is trained as a speaker-verification model — a system whose only job is to decide whether two audio clips are spoken by the same person or different people, trained on datasets like VoxCeleb1 that provide thousands of speakers' worth of audio without needing word-level transcripts. Solving that verification task forces the model to learn a compact numerical representation (an embedding) that captures the character of a voice, independent of what's being said. Because this embedding space generalizes well, it can produce a usable representation of a voice the encoder has never encountered before, from just a few seconds of that voice speaking.
The synthesizer is a Tacotron-style sequence-to-sequence model that takes text and that speaker embedding together, and generates a mel spectrogram — a visual, frequency-based representation of speech — conditioned on both. Trained on datasets like LibriSpeech and VCTK, it learns to read arbitrary text in a way modulated by whatever embedding it's given, which is what lets a single trained synthesizer produce speech in many different voices rather than needing a separate model per speaker.
The vocoder takes that mel spectrogram and converts it into an actual playable waveform. The vocoder here is based on WaveRNN, a neural architecture specifically chosen for its ability to generate audio fast enough for real-time use, rather than the slower, higher-latency vocoders common at the time. Because these three stages are trained largely independently — the encoder on a speaker-verification dataset, the synthesizer on transcribed multi-speaker audio, and the vocoder on the resulting spectrograms — a new voice can be cloned at inference time just by running a short sample through the already-trained encoder, without retraining the synthesizer or vocoder at all.
Getting Started with Real-Time-Voice-Cloning Locally
- Clone the repository and prepare the environment. The current repository uses uv to manage Python and its dependencies. Its configuration requires Python 3.9 and PyTorch 1.10 for both CPU and CUDA options. uv creates the appropriate environment when you run a command.
- Install ffmpeg and uv. ffmpeg is needed to read audio files. After installing uv, run the appropriate command below from the repository root; uv will handle the Python dependencies.
- Use the pretrained models. The encoder, synthesizer, and vocoder weights download automatically on first run, so you do not need to train the models yourself. If the automatic download fails, use the manual links in the official README.
- Run the toolbox for an interactive session.
demo_toolbox.pyopens a GUI where you can record or load a reference voice, generate speech from typed text, and inspect the resulting embedding visually. - Or use the CLI for a scriptable path.
demo_cli.pyruns the same encoder-synthesizer-vocoder pipeline from the command line, better suited to integrating into a larger script than the interactive toolbox. - Use a GPU if you plan to retrain anything. The toolbox alone can run on more modest hardware (an experimental low-memory mode supports GPUs with as little as ~2GB VRAM), but retraining any of the three models in earnest benefits significantly from a dedicated GPU.
| Usage | CPU | NVIDIA GPU |
|---|---|---|
| Graphical toolbox | uv run --extra cpu demo_toolbox.py | uv run --extra cuda demo_toolbox.py |
| Command-line example | uv run --extra cpu demo_cli.py | uv run --extra cuda demo_cli.py |
Tips for Better Results with Real-Time-Voice-Cloning
- Set expectations before you start. Even the original author has been candid that output quality lags behind modern commercial and open-source alternatives — approach this as a way to understand how voice cloning works, not to produce publication-ready audio.
- Use clean, single-speaker reference clips. Since the encoder was trained on relatively clean speaker-verification data, background noise or overlapping speech in your reference sample degrades the resulting embedding more noticeably than it would in newer, more robust systems.
- Inspect the embedding visualization in the toolbox before generating. The GUI's embedding plot can reveal when a reference clip produced a poor or unusual representation, which is faster to catch there than after listening to garbled synthesizer output.
- Start with the pretrained models before attempting to retrain. Downloading LibriSpeech's train-clean-100 subset alone is enough to experiment with the toolbox — save the full VoxCeleb1 and VCTK downloads for when you actually intend to retrain a model from scratch.
- Check MockingBird if your target language isn't English. Rather than trying to adapt the original English-trained models to another language yourself, the existing Mandarin-focused fork is a more direct starting point for that specific case.
Real-Time-Voice-Cloning vs. Newer Open-Source Voice Cloning
| Real-Time-Voice-Cloning (SV2TTS) | Newer open-source projects (e.g., Coqui XTTS-v2, Chatterbox) | |
|---|---|---|
| Released | 2019 | 2023–2025 |
| Maintenance | Older architecture; later dependency, installation, and compatibility updates | Actively maintained |
| Architecture | Separate encoder, synthesizer, vocoder (Tacotron + WaveRNN) | Often single end-to-end or more integrated architectures |
| Zero-shot quality | Historically significant, but dated by current standards | Meaningfully more natural and stable |
| Best use today | Learning how voice cloning works, hands-on | Actual self-hosted production or hobbyist use |
Frequently Asked Questions
What is SV2TTS?
SV2TTS (Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis) is the three-stage technique Real-Time-Voice-Cloning is built around, originally developed by a Google research team in 2018.
Is Real-Time-Voice-Cloning still maintained?
The model architecture is older, but the repository has received later dependency, installation, and compatibility updates. The author also cautions that its audio quality trails newer options, so evaluate whether it fits a production project.
How much audio does Real-Time-Voice-Cloning need to clone a voice?
As little as a few seconds, per the project's original tagline — though real-world results depend heavily on the cleanliness of the reference sample and generally lag behind newer voice cloning systems in overall quality.
Is Real-Time-Voice-Cloning free to use?
Yes. It runs on your own hardware, with no account, API key, or subscription required — though you'll need a reasonably capable computer, ideally with a GPU, to run it comfortably.
What's the relationship between this project and Resemble AI?
Corentin Jemine, the project's creator, later went to work at Resemble AI, and Resemblyzer — a standalone package derived from this project's speaker encoder — is now maintained under Resemble AI's GitHub organization.
Generate New Speech from a Reference Voice in Echora
For a similar voice-cloning workflow in your browser, upload a reference recording to Echora’s Voice Clone, enter new text, and generate speech that resembles the reference voice. You can then preview and download the result. Use only your own voice or a voice you have explicit permission to clone.