MockingBird: A Community Fork That Taught an English Voice Cloner to Speak Mandarin
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Real-Time-Voice-Cloning, the project MockingBird is built on, could clone a voice from a few seconds of audio — but only in English. Developer babysor forked it specifically to close that gap, retraining and extending it for Mandarin Chinese rather than starting from scratch. That community-driven origin story is exactly why MockingBird became one of the most forked and starred Chinese voice-cloning repositories on GitHub: it solved a real, specific gap for Chinese-speaking developers rather than being a general research release that happened to support the language.
What Is MockingBird?
MockingBird is an open-source voice cloning and real-time speech synthesis toolkit, built as a fork of the SV2TTS-based Real-Time-Voice-Cloning project and re-engineered specifically for Mandarin Chinese. It clones a target voice from as little as 5 seconds of reference audio and generates arbitrary new speech in that voice in real time, with training and testing conducted across several major Chinese speech datasets.
Functional Advantages of MockingBird
- 5-second voice cloning. A short reference clip is enough to capture a target voice, thanks to reusing a pretrained speaker-verification encoder rather than training voice recognition from scratch for each new speaker.
- Real-time generation, with an inference engine specifically optimized for speed rather than only batch-style offline processing.
- Purpose-built for Mandarin, trained and validated across multiple Chinese speech datasets — aidatatang_200zh, MAGICDATA, AISHELL-3, and data_aishell among them — rather than treating Chinese as an afterthought bolted onto an English-first model.
- Genuinely cross-platform, running on Windows, Linux, and even Apple Silicon Macs.
- Two ways to use it: a graphical toolbox for interactive, local experimentation, and a webserver mode for remote calls — covering both hands-on testing and integration into a larger application.
- A large, active Chinese developer community, reflected in its GitHub fork and star counts, which also means more community troubleshooting resources and shared experience than a typical academic release tends to have.
How the Transfer-Learning Trick Actually Works
This is the specific engineering decision that made a Chinese-language voice cloning project achievable without massive new training resources, and it's worth understanding. MockingBird's pipeline, inherited from the SV2TTS approach it's built on, splits into three components: an encoder that learns to represent a speaker's voice characteristics (originally trained for a speaker-verification task, which generalizes well to identifying "what makes this voice sound like this voice" even for speakers never seen during training), a synthesizer that converts text into a mel-spectrogram conditioned on that speaker representation, and a vocoder that turns the spectrogram into an actual waveform.
MockingBird's approach reuses the pretrained encoder and vocoder from the original English project largely as-is, and focuses new training specifically on the synthesizer — the component responsible for actually mapping Chinese text to speech content. Because the encoder's speaker-identification capability and the vocoder's waveform-generation capability don't need to be relearned from zero, adding Mandarin support became a considerably more tractable project than training an entirely new three-stage pipeline from scratch, which is directly why the project's own documentation describes getting "easy and awesome" results from retraining just one of the three components.
Getting Started with MockingBird
Setup starts with creating a dedicated Python 3.9 conda environment, cloning the repository, and installing dependencies via pip — including webrtcvad-wheels specifically, alongside PyTorch and its related packages. Once your environment is ready, prepare a 5-to-30-second audio sample of your target voice, then launch the provided graphical toolbox to load your sample, type in the text you want spoken, and generate cloned speech directly through the interface. For integration into another application rather than interactive use, the webserver mode exposes the same capability for remote calling instead.
Tips for Better Results
- Use a clean recording somewhere in the 5-to-30-second range. While 5 seconds is the headline minimum, a slightly longer, well-recorded sample tends to produce more stable results than pushing right up against the minimum every time.
- Expect to troubleshoot dependency versions on newer systems. The project's documented testing was done against PyTorch 1.9.0 and specific GPU configurations from around 2021 — on a modern setup, you may need to adjust package versions manually rather than assuming every dependency installs cleanly out of the box.
- Use the GUI toolbox first before building a custom integration. It's the fastest way to confirm your environment and a given voice sample actually work before writing any code around the webserver mode.
- Lean on the Chinese developer community for setup help. Given how widely forked and discussed this project is within that community, a specific error message you hit has likely already been solved and documented by someone else.
MockingBird vs. Other Chinese-Focused Cloning Tools
| MockingBird | GPT-SoVITS | XTTS-v2 | |
|---|---|---|---|
| Minimum clip for cloning | 5 seconds | 5 seconds (zero-shot) | ~6 seconds |
| Core architecture | SV2TTS-style encoder/synthesizer/vocoder | GPT + SoVITS (VITS-based) hybrid | GPT-2-style autoregressive + VQ-VAE |
| Chinese-specific origin | Yes, forked specifically to add Mandarin | Strong Chinese/Japanese reputation | General multilingual, not Chinese-specific |
| Built-in data prep tools | No | Yes (vocal separation, ASR, auto-slicing) | No |
| Interface options | GUI toolbox + webserver | Full WebUI | Python API |
| Community activity level | Large historical community, less active core development | Actively maintained, frequent version updates | Community-maintained forks post-Coqui shutdown |
MockingBird's specific niche is being the original, community-driven answer to "I want the Real-Time-Voice-Cloning experience, but in Mandarin" — newer projects like GPT-SoVITS have since built more actively maintained, feature-rich alternatives, but MockingBird's popularity and documentation base remain a genuinely useful starting point, especially for anyone learning how SV2TTS-style architectures work under the hood.
Frequently Asked Questions
How much audio does MockingBird need to clone a voice?
As little as 5 seconds, with a recommended range of 5 to 30 seconds for more stable results.
Why was MockingBird created instead of just using the original English project?
The original Real-Time-Voice-Cloning project only supported English. MockingBird was forked specifically to retrain and extend it for Mandarin Chinese, using multiple Chinese speech datasets for training and validation.
Does MockingBird require training a new voice encoder for Chinese?
No — that's the specific efficiency behind the project. It reuses the pretrained encoder and vocoder from the original English model and focuses new training on the synthesizer component, which is what makes adding Chinese support achievable without retraining the entire pipeline.
What platforms does MockingBird run on?
Windows, Linux, and Apple Silicon Macs are all documented as supported.
Is MockingBird actively maintained?
Its core development pace has slowed since its most active period, and its documented testing reflects software versions from around 2021 — expect to troubleshoot some dependency versions on a modern system, though its large community means most common issues already have documented solutions.
Create a Similar Voice from Reference Audio
For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.