EchoraEchora
Back to Models

GPT-SoVITS: A Full Voice-Cloning Studio, Not Just a Model

September 2, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Plenty of open-source cloning models hand you a checkpoint and a Python script and call it done. GPT-SoVITS hands you the entire workflow instead — vocal separation, automatic transcription, dataset slicing, training, and inference, all inside one browser-based interface — built around a genuinely low data requirement: a usable clone from 5 seconds of audio, and a properly fine-tuned one from just a minute.

What Is GPT-SoVITS?

GPT-SoVITS is an open-source, MIT-licensed voice cloning and text-to-speech project created by developer RVC-Boss, combining GPT-based semantic text modeling with SoVITS — a VITS-based acoustic and voice conversion system — into a single pipeline. The GPT component handles converting text into semantic tokens; the SoVITS component then converts those tokens, guided by the reference voice, into the final waveform. That division of labor is what lets the project support both instant zero-shot generation and deeper few-shot fine-tuning from the same underlying architecture.

It supports cross-lingual generation across Chinese, English, Japanese, Korean, and Cantonese, and has built a particularly strong reputation specifically for Chinese and Japanese output quality.

Two Tiers of Cloning: 5 Seconds and 1 Minute

This is the headline capability, and it's worth being precise about what each tier actually gets you. Zero-shot mode takes a 5-second reference clip and generates speech in that voice immediately, no training step required — developer testing has described this tier as producing roughly 80-95% similarity to the target voice, which is usable for quick tests or lower-stakes content. Few-shot fine-tuning takes that further: with about a minute of target voice data, the model can be fine-tuned to noticeably improve similarity and naturalness beyond the zero-shot baseline, closing much of the remaining gap toward the reference speaker's actual voice.

The Version History: v1 Through v4

GPT-SoVITS has gone through several distinct architectural generations since its initial release, and knowing which one fits your situation matters more than defaulting to "the newest":

  • v1 and v2 share the same fundamental architecture — direct waveform generation through a VQ-VAE decoder without a separate vocoder stage. v2 extended this with broader language support and more pretraining data.
  • v2Pro / v2ProPlus refined stability and compatibility further while keeping v2's hardware footprint and inference speed, and in the project's own comparisons, actually outperforms v4 in some respects while staying cheaper to run.
  • v3 replaced the core acoustic architecture with a shortcut Conditional Flow Matching Diffusion Transformer (shortcut-CFM-DiT), producing noticeably better timbre similarity to the reference audio, fewer GPT-side repetition or omission errors, and richer emotional expression — trained on an expanded, quality-filtered dataset of roughly 7,000 hours. Because generation isn't fully end-to-end in this version, it relies on the open-source BigVGAN v2 vocoder to convert its mel-spectrogram output into 24kHz audio, which can introduce a metallic artifact from non-integer-multiple upsampling.
  • v4 fixes that specific metallic-sound issue and natively outputs 48kHz audio instead of 24kHz, avoiding the slightly muffled quality that could result from upsampling a 24kHz signal. The project's own author describes v4 as a practical drop-in replacement for v3.

The practical decision point: v1, v2, and v2Pro tend to perform better on lower-quality or noisy training data, because their architecture is more influenced by the overall average of the training set. v3 and v4 perform better specifically on clean, high-quality reference audio, leaning more heavily toward faithfully reproducing that exact reference rather than averaging across a dataset — but they can actually underperform v1/v2 if your source recordings aren't clean to begin with.

The WebUI: Built for People Who Aren't Training Pipelines by Hand

GPT-SoVITS's interface bundles the entire dataset preparation workflow, not just inference. It includes UVR5 for separating vocals from background music or noise, automatic speech recognition for transcription (Chinese ASR through dedicated models, English and Japanese through Faster Whisper Large V3), and automatic audio slicing to break longer recordings into usable training segments. For smaller datasets in the 5-10 minute range, a "one-click triple action" button chains several of these preprocessing steps together automatically — a meaningful accessibility advantage if you're not comfortable scripting a data pipeline from scratch.

Getting Started with GPT-SoVITS

  1. Use the Windows integrated package if you're on Windows. A prebuilt package is available for direct download — double-click go-webui.bat and the full WebUI launches without a manual environment setup.
  2. Set up manually on Linux/Mac, or if you want more control. Create a Python 3.10 virtual environment (via uv or conda), clone the repository, and install dependencies with pip install -r requirements.txt.
  3. Download the pretrained models matching your target version. Each version (v2, v2Pro, v3, v4) has its own specific checkpoint files hosted on Hugging Face — make sure the pretrained model version matches the training version you select in the WebUI, or generation will fail or degrade.
  4. Prepare your reference audio with the built-in tools. Run UVR5 first if your source has background noise or music, then use the automatic slicing and ASR annotation tools rather than manually transcribing and cutting clips yourself.
  5. Pick your version based on source audio quality. Use v1/v2/v2Pro for noisier or more average-quality recordings; use v3 or v4 specifically when your reference audio is clean studio-quality and you want maximum fidelity to that exact voice.
  6. Watch for the 16-series GPU issue. Older NVIDIA 16-series cards lack tensor cores and can throw a CUDA configuration error during training; the project's config.py auto-detects this and forces FP32 precision as a workaround.

Tips for Better Results

  • Start with zero-shot before committing to fine-tuning. A quick 5-second test tells you whether the base model's zero-shot similarity is already good enough for your use case before investing time in gathering and preparing a full minute of training audio.
  • Match version to data quality honestly, not by recency. Defaulting to v4 because it's the newest can actually hurt output quality if your source recordings are noisy — v2Pro is frequently the more practical choice for less-than-pristine audio.
  • Clean your dataset before training, not after. Use UVR5 to strip background noise and music first; both v3 and v4 in particular lean heavily on the fidelity of what you feed them.
  • Lean into its Chinese and Japanese strength deliberately. If your project's primary language is Chinese or Japanese, GPT-SoVITS is one of the stronger open-source options specifically for those languages, more so than for English-only content where other models may have an edge.
  • Prefer v4 over v3 unless you have a specific reason not to. Since v4 fixes v3's known metallic-artifact issue while outputting higher native sample-rate audio, the author's own guidance to treat it as a drop-in replacement is a reasonable default unless you're specifically troubleshooting a v3-based setup.

GPT-SoVITS vs. Other Few-Shot Cloning Models

GPT-SoVITSXTTS-v2Fish Speech
Minimum data for a usable clone5 seconds (zero-shot)~6 seconds~10–30 seconds
Fine-tuning supportYes, ~1 minute of dataYes, via community scriptsYes, native
Strongest languagesChinese, JapaneseBroad 17-language coverageEnglish, Chinese, Japanese
Built-in data prep toolsYes (vocal separation, ASR, auto-slicing)NoNo
LicenseMITCoqui Public Model License (non-commercial)Research/non-commercial-leaning

GPT-SoVITS's specific niche is the combination of a genuinely low data requirement with an all-in-one WebUI that handles data preparation as well as training — for Chinese and Japanese content especially, that combination is hard to find bundled together elsewhere.

Frequently Asked Questions

How much audio do I actually need to clone a voice with GPT-SoVITS?

As little as 5 seconds for an immediate zero-shot clone, or about 1 minute of reference audio for a fine-tuned model with noticeably better similarity and naturalness.

Which version should I use, v2, v3, or v4?

Use v2 or v2Pro if your training audio is average quality or noisy; use v3 or v4 if your reference recordings are clean, studio-quality audio, since those versions lean more heavily on faithfully matching the exact reference. Between v3 and v4 specifically, v4 fixes a known metallic-artifact issue and outputs higher native audio quality, so it's the more practical default.

Is GPT-SoVITS good for languages other than Chinese?

Yes — it supports English, Japanese, Korean, and Cantonese alongside Chinese — but Chinese and Japanese are where it has built the strongest reputation.

Is GPT-SoVITS free for commercial use?

Yes. It's released under the MIT license, one of the more permissive options among open-source voice cloning projects.

Do I need to prepare my own training dataset manually?

Not necessarily. The WebUI includes vocal separation, automatic transcription, and audio slicing tools, with a one-click option for smaller datasets that chains several preprocessing steps together automatically.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.

Create a Voice Clone →