EchoraEchora
Back to Models

MaskGCT: Solving Two Opposite TTS Problems With One Training Trick

September 12, 2026
•
6 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Autoregressive TTS models model duration implicitly, which tends to hurt robustness and gives you no real control over how long the output actually runs. Non-autoregressive models fix the robustness issue, but historically needed explicit text-speech alignment supervision and phone-level duration prediction during training — extra machinery that can itself compromise naturalness. MaskGCT, developed by researchers at CUHK-Shenzhen and released through the Amphion toolkit, was built to avoid both problems at once, using a training approach borrowed from a different corner of generative modeling entirely: mask, then predict.

What Is MaskGCT?

MaskGCT (Masked Generative Codec Transformer) is a fully non-autoregressive, open-source zero-shot TTS model, released in October 2024 and later published at ICLR 2025. It eliminates the need for explicit alignment information between text and speech during training, and does away with phone-level duration prediction entirely — a genuine architectural departure from typical non-autoregressive TTS systems, not just a tuning difference. It's distributed through Amphion, an open-source toolkit for audio, music, and speech generation research.

Functional Advantages of MaskGCT

  • No alignment supervision required during training, removing a whole category of preprocessing and potential failure point that most non-autoregressive TTS systems depend on.
  • No phone-level duration prediction, sidestepping a mechanism that can itself introduce unnaturalness into generated speech.
  • Explicit control over total output duration, letting you specify exactly how long a generated clip should run — a genuine capability advantage over autoregressive models, which model duration only implicitly.
  • Zero-shot voice cloning from just 5 seconds of reference audio, requiring no additional training for the target speaker.
  • Fast, parallel generation, needing only 25 to 50 inference steps for the text-to-semantic stage rather than generating tokens one at a time in sequence.
  • Demonstrated versatility beyond core TTS, with the same underlying system explored for cross-lingual dubbing, voice conversion, emotion control, and speech content editing.

How Mask-and-Predict Generation Actually Works

MaskGCT's training approach is conceptually similar to masked language modeling in text (think BERT) or masked generative modeling in images (think MaskGIT), applied here to discrete speech tokens. During training, the model learns to predict masked semantic or acoustic tokens given surrounding context, conditions, and prompts. During inference, rather than generating tokens sequentially, it generates a full sequence of tokens at a specified length through iterative parallel decoding — filling in the masked positions across several passes rather than one token at a time.

The actual pipeline runs in two stages. The text-to-semantic (T2S) stage predicts semantic tokens from input text and prompt semantic tokens; these semantic tokens are extracted using a VQ-VAE trained specifically to quantize embeddings from a speech self-supervised learning model, a deliberate choice over the k-means clustering common in earlier work, since it minimizes information loss even with a single codebook. The semantic-to-acoustic (S2A) stage then predicts acoustic tokens — drawn from an RVQ-based speech codec — conditioned on those semantic tokens plus additional acoustic prompts, using layer-wise generation across the codec's multiple token layers to preserve fine acoustic detail.

Benchmark Results

Trained on 100,000 hours of in-the-wild speech from the Emilia dataset (split evenly between English and Chinese), MaskGCT reported outperforming the prior state-of-the-art zero-shot TTS systems it was compared against on quality, similarity, and intelligibility, reaching what its authors describe as human-level results on these metrics. In a direct comparison against a strong autoregressive-plus-SoundStorm baseline, MaskGCT showed measurable naturalness improvements (CMOS gains of +0.12 on LibriSpeech test-clean, +0.08 on SeedTTS test-en, and +0.37 on SeedTTS test-zh) alongside better similarity and robustness. The robustness advantage was particularly pronounced on genuinely hard cases — repeated words and tongue twisters, the classic triggers for TTS hallucination and word-skipping — which is a meaningfully practical result beyond just topline benchmark scores.

Getting Started with MaskGCT

  1. Clone the Amphion repository. git clone https://github.com/open-mmlab/Amphion.git gets you the full toolkit, of which MaskGCT is one component.
  2. Set up the environment using the provided script. bash ./models/tts/maskgct/env.sh handles the MaskGCT-specific environment setup within the broader Amphion repository.
  3. Download the pretrained checkpoints from Hugging Face. Use hf_hub_download("amphion/MaskGCT", filename="semantic_codec/model.safetensors") as the pattern, repeating for the other required components — the acoustic codec encoder and decoder, the T2S model, and the S2A models (both the 1-layer and full variants).
  4. Use the MaskGCT_Inference_Pipeline class for generation. Instantiate it with your loaded semantic model, semantic codec, codec encoder/decoder, T2S model, and S2A models, then call .maskgct_inference() with your prompt audio path, prompt text, target text, and source/target language codes.
  5. Set target_len explicitly when you need a precise output duration. Passing a specific value in seconds gives you direct control over generated length; leaving it as None falls back to a simple duration-estimation rule instead.
  6. Try the Jupyter notebook, Gradio interface, or Hugging Face Space demo if you'd rather test the model interactively before scripting a custom pipeline.

Tips for Better Results

  • Use target_len deliberately whenever timing matters. This is one of MaskGCT's genuine structural advantages over autoregressive models — take advantage of it for any use case (video narration, dubbing to a fixed slot) where exact duration is a real constraint rather than a nice-to-have.
  • Test on genuinely hard text, not just clean sentences. Since MaskGCT's robustness advantage shows up most clearly on repeated words and tongue-twister-style content, that's also the most informative way to evaluate it against alternatives for your own use case.
  • Keep your reference clip clean for the full 5-second minimum. Since cloning quality depends on the semantic and acoustic tokens extracted from that reference, a clear recording at or above that length will outperform a shorter or noisier sample.

MaskGCT vs. Other Non-Autoregressive TTS Models

MaskGCTF5-TTSE2 TTS
Core approachMask-and-predict, two-stage semantic/acoustic tokensDiffusion Transformer + flow matchingFlow matching, flat U-Net Transformer
Alignment supervision neededNoNoNo
Explicit duration controlYes, via target_len parameterNo dedicated duration modelNo dedicated duration model
Reference clip for cloning~5 seconds~5–15 secondsSimilar
Languages6 (English, Chinese, Korean, Japanese, French, German)Primarily English-focusedEnglish-focused

MaskGCT's specific niche is combining alignment-free, duration-prediction-free training with genuine explicit control over output length — a combination that directly addresses the specific weaknesses of both the autoregressive and typical non-autoregressive camps it positions itself between.

Frequently Asked Questions

What does "masked generative" mean in MaskGCT's name?

It refers to the model's mask-and-predict training paradigm, conceptually similar to masked language modeling in text or masked generative image models — the model learns to predict masked tokens from surrounding context, then generates full sequences through iterative parallel decoding at inference time.

Can I control exactly how long the generated speech is?

Yes, through the target_len parameter, which lets you specify an exact output duration — a capability autoregressive models don't offer directly, since they model duration only implicitly.

How much reference audio does MaskGCT need for voice cloning?

As little as 5 seconds, with no additional training required for the target speaker.

What languages does MaskGCT support?

Six: English, Chinese, Korean, Japanese, French, and German.

Is MaskGCT ready for commercial or production use?

The official project page describes it as being for research demonstration purposes — test thoroughly on your own content and review current license terms before considering any production deployment.

Create Similar-Sounding Speech from a Reference Voice

For a browser-based workflow with a similar goal, use Echora's Voice Clone. Upload a recording of your own voice or a voice you have explicit permission to use, enter new text, and generate speech that resembles the reference.

Create a Voice Clone →