EchoraEchora
Back to Models

GLM-TTS: Zhipu's Production-Grade Answer to "Good Enough" Chinese TTS

September 4, 2026
•
7 min read

A lot of open-source TTS models are research demos wearing a GitHub README. GLM-TTS, released by Zhipu AI in collaboration with Tsinghua University, was built with a different target from the start: production deployment. It clones a voice from a 3-second clip, controls emotion and pronunciation with genuine precision, and was trained on a comparatively modest 100,000 hours of data while still landing state-of-the-art results on multiple open benchmarks — efficiency being as much the point as raw quality.

What Is GLM-TTS?

GLM-TTS is an open-source, industrial-grade text-to-speech system developed by Zhipu AI and Tsinghua University researchers, built around a two-stage architecture: a text-to-token autoregressive language model, followed by a token-to-waveform flow-matching model. That combination — language model for content and structure, flow matching for acoustic realization — is a similar broad pattern to other modern TTS systems, but GLM-TTS distinguishes itself through a specific set of production-focused engineering choices layered on top: an optimized speech tokenizer, a reinforcement-learning training framework tuned for multiple quality dimensions at once, and lightweight tools for deploying custom voices without retraining the whole model.

Under the Hood: LLM Meets Flow Matching

The first stage is a Llama-architecture-based autoregressive language model that converts input text, combined with a speaker embedding, into a sequence of speech tokens. That speaker embedding is extracted by a CAMPPlus model in the system's frontend — feed it any 3 to 10 second reference clip, and GLM-TTS reproduces that voice's timbre and prosody immediately, with no per-speaker fine-tuning required.

The second stage takes those tokens and turns them into actual audio using a Diffusion Transformer with continuous flow matching, converting the token sequence into high-quality mel-spectrograms before a vocoder renders the final waveform. GLM-TTS's speech tokenizer itself was specifically optimized with fundamental frequency (pitch) constraints built in, which is part of why pitch accuracy and overall speech quality hold up despite the comparatively modest 100,000-hour training scale.

GRPO Reinforcement Learning: Optimizing Three Things at Once

This is GLM-TTS's most technically distinctive contribution, and it's worth understanding why it matters. Most TTS training optimizes for a single dominant objective — usually intelligibility or speaker similarity — and treats other qualities like emotional expressiveness as secondary. GLM-TTS instead uses a GRPO-based (Group Relative Policy Optimization) multi-reward reinforcement learning framework that jointly optimizes pronunciation accuracy, speaker similarity, and expressive prosody together, rather than sequentially or in isolation. That joint optimization is the specific mechanism behind GLM-TTS avoiding the flat, "mechanical" quality that plagues a lot of TTS output that technically pronounces words correctly but doesn't sound like anyone actually talking.

Emotional and Paralinguistic Control

Built on that same RL framework, GLM-TTS can generate specific emotional deliveries — happy, sad, angry — alongside natural paralinguistic sounds like laughter and breathing, directly as part of generation rather than through separate post-processing. This is reinforcement-learning-driven expressiveness rather than a fixed emotion parameter, which is part of why the resulting delivery tends to sound more organically tied to the actual content than a simple style tag might produce.

Phoneme-In: Precise Pronunciation Without Sacrificing Flexibility

GLM-TTS supports a hybrid phoneme-and-text input scheme, referred to as "Phoneme-in," letting you mix explicit phonetic notation directly into otherwise normal text input. This is specifically aimed at the polyphone and rare-word problem that trips up many TTS systems — where a specific character or word has ambiguous pronunciation options — giving you a direct, surgical way to lock in the correct reading for just the problematic word or phrase, without having to phonetically spell out an entire sentence.

LoRA-Based Voice Customization

For deployment scenarios that need a specific, consistent custom voice beyond what zero-shot cloning reliably delivers, GLM-TTS supports parameter-efficient customization through LoRA (Low-Rank Adaptation) — a lightweight fine-tuning approach that adapts the model to a target voice without the cost and complexity of retraining the full network. This is the practical middle ground between "good enough" zero-shot cloning and a full custom training pipeline.

Performance and Streaming

Thanks to its flow-matching-based decoder, GLM-TTS supports streaming inference natively, and on an RTX 4090, the system runs at roughly 3 to 5 times real-time speed — fast enough for live assistants, game character dialogue, and streaming commentary use cases where waiting for a full clip to finish generating isn't an option.

Getting Started with GLM-TTS

  1. Check your hardware first. GLM-TTS is built for Python 3.10–3.12 and expects an NVIDIA GPU with at least 8GB of VRAM plus a working CUDA toolkit; budget roughly 9GB of disk space for the model weights. CPU inference is possible but dramatically slower — expect 5 to 15 minutes per generation rather than the few seconds a GPU delivers.
  2. Download the model weights. Use huggingface-cli download zai-org/GLM-TTS --local-dir ckpt, or pull from ModelScope with modelscope download --model ZhipuAI/GLM-TTS --local-dir ckpt if that's a faster source for your location.
  3. Install dependencies. Set up a Python 3.10–3.12 environment and install from the project's requirements.txt before running any inference scripts.
  4. Run a basic generation. python glmtts_inference.py --data=example_zh --exp_name=_test --use_cache walks through the documented example workflow using the provided sample data.
  5. Add the --phoneme flag for precise pronunciation control. When you need to lock in a specific reading for a polyphone or rare word, enabling phoneme capabilities through this flag gives you the hybrid phoneme-text input path rather than relying on the model's default text-only inference.
  6. Try the Gradio app for quick testing. Upload a reference clip, type your text, adjust generation parameters like temperature and top_p, and listen to the output directly, rather than scripting every test generation manually.

Tips for Better Results

  • Use a genuinely clean 3–10 second reference clip. Since the CAMPPlus speaker embedding is doing the heavy lifting for zero-shot cloning, background noise or a poor recording in your reference clip will degrade timbre reproduction more than it would with some other architectures.
  • Reach for Phoneme-in specifically for polyphones and rare words, not entire scripts. It's designed as a targeted correction tool — mixing in phonetic notation only where pronunciation is genuinely ambiguous keeps your workflow simpler than phoneticizing full sentences unnecessarily.
  • Use LoRA customization when zero-shot isn't consistent enough. If a specific voice needs to sound reliable across a large volume of content rather than just a single clip, the LoRA path is the more practical option than repeatedly hoping zero-shot cloning nails it each time.
  • Match your GPU to your latency needs. The documented 3–5x real-time figure is specific to an RTX 4090 — if you're deploying on more modest hardware, benchmark your actual throughput before committing to a real-time or streaming use case.
  • Don't rely on CPU inference for anything time-sensitive. Given the 5-to-15-minute-per-generation gap between CPU and GPU performance, CPU inference is realistically limited to offline, non-interactive use cases rather than any live application.

GLM-TTS vs. Other Strong Chinese-Capable Open Models

GLM-TTSCosyVoice3GPT-SoVITS
Backing organizationZhipu AI + Tsinghua UniversityAlibaba (FunAudioLLM)RVC-Boss (independent)
Zero-shot cloningYes, ~3–10 second referenceYes, ~3–10 second referenceYes, 5 seconds
Training approachGRPO multi-reward RL (pronunciation + similarity + prosody jointly)Supervised trainingSupervised + optional fine-tuning
Emotion controlRL-driven (happy/sad/angry, laughter, breathing)Instruction-basedNot a dedicated feature
Pronunciation correctionPhoneme-in hybrid inputPronunciation inpaintingManual annotation
StreamingYes, nativeYesNo
LicenseApache 2.0 (code) / MIT (weights)Open, hosted API availableMIT

GLM-TTS's specific contribution is treating pronunciation, speaker similarity, and emotional prosody as objectives to optimize together rather than in sequence — a training-methodology difference that's less visible than a headline feature, but is exactly the kind of engineering choice that separates a genuinely production-ready system from a strong research demo.

Frequently Asked Questions

Who developed GLM-TTS?

GLM-TTS was developed by Zhipu AI in collaboration with Tsinghua University, released as open source in December 2025.

How much reference audio does GLM-TTS need for voice cloning?

Roughly 3 to 10 seconds, with a speaker embedding extracted via a CAMPPlus model enabling zero-shot cloning with no per-speaker fine-tuning required.

What is GRPO in the context of GLM-TTS?

It's Group Relative Policy Optimization, a reinforcement learning framework GLM-TTS uses to jointly optimize pronunciation accuracy, speaker similarity, and expressive prosody together during training, rather than treating them as separate, sequential objectives.

Is GLM-TTS free for commercial use?

Yes. GLM-TTS's code is released under the Apache 2.0 license and its model weights under the MIT license, both of which permit commercial use.

What hardware do I need to run GLM-TTS?

An NVIDIA GPU with at least 8GB of VRAM, CUDA toolkit installed, Python 3.10–3.12, and roughly 9GB of disk space for the model weights. CPU inference works but is significantly slower.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.

Create a Voice Clone →