EchoraEchora
Back to Models

VoxCPM: Skipping the Tokenization Step Most TTS Models Rely On

September 10, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Nearly every modern TTS architecture — from VALL-E's codec tokens to CosyVoice's semantic tokens to GPT-SoVITS's discrete codebooks — starts by compressing speech into discrete units before a model learns to predict them. VoxCPM, from OpenBMB, takes a different route entirely: it models speech directly in continuous space, skipping discrete tokenization altogether. That architectural choice is the foundation for both of the model's flagship capabilities — context-aware prosody and genuinely convincing zero-shot voice cloning — in a bilingual Chinese-English model fast enough to stream on a single consumer GPU.

What Is VoxCPM?

VoxCPM is a tokenizer-free text-to-speech system built by OpenBMB on top of the MiniCPM-4 language model backbone, released under the Apache-2.0 license. Rather than converting speech into discrete tokens the way most comparable systems do, it uses an end-to-end diffusion autoregressive architecture to generate continuous speech representations directly from text, trained on more than 1.8 million hours of bilingual Chinese-English audio.

Functional Advantages of VoxCPM

  • Tokenizer-free continuous modeling, avoiding the information loss that discrete tokenization can introduce, while still preserving generation stability through a semi-discrete bottleneck.
  • Context-aware speech generation — the model infers appropriate prosody and emotional tone directly from the meaning of the input text, with no manual style tags required.
  • State-of-the-art benchmark results among open-source systems, outperforming strong competitors including IndexTTS2 and CosyVoice2 on the SEED-TTS-EVAL benchmark.
  • Genuinely low latency, achieving a real-time factor as low as 0.17 on a consumer-grade NVIDIA RTX 4090 — meaning it generates audio roughly six times faster than the resulting speech takes to play.
  • True-to-life zero-shot voice cloning, with high speaker similarity scores across both supported languages.
  • Openly licensed with fine-tuning support, released under Apache-2.0 with both supervised fine-tuning (SFT) and LoRA adaptation documented for customizing the model further.

How Tokenizer-Free Generation Actually Works

This is the architectural decision that sets VoxCPM apart, and it's worth understanding why it matters. Discrete-token TTS systems compress speech into a fixed vocabulary of codes — efficient for language-model-style prediction, but lossy by construction, since continuous acoustic detail gets rounded into a limited set of discrete symbols. VoxCPM instead combines hierarchical language modeling, Finite Scalar Quantization (FSQ), and local Diffusion Transformers to generate continuous speech representations directly, sidestepping that information bottleneck while still keeping generation stable enough to train reliably.

The FSQ component specifically produces what the project describes as a structured semi-discrete representation — a middle ground that implicitly separates high-level semantic content from fine-grained acoustic detail without forcing either into a fully discrete code. That implicit disentanglement is what enables the context-aware generation capability: because semantic meaning and acoustic texture aren't tangled together in a single representation, the model can infer prosody and emotional tone from what the text actually means, rather than requiring you to specify delivery style as a separate instruction.

Benchmark Results

On the SEED-TTS-EVAL benchmark, VoxCPM-0.5B reported a 1.85% word error rate in English and a 0.93% character error rate in Chinese — both ahead of strong competitors including IndexTTS2 and CosyVoice2 — while maintaining speaker similarity scores of 72.9% in English and 77.2% in Chinese. A variant trained on the smaller, fully public Emilia dataset (VoxCPM-Emilia) still posted competitive results (2.34% English WER, 1.11% Chinese CER), which the project points to as evidence of the architecture's data efficiency rather than results dependent purely on training-set scale. In subjective listening evaluations, VoxCPM led on English speaker similarity and scored well on naturalness, while for Chinese it trailed IndexTTS2 slightly on naturalness but edged ahead on speaker similarity — suggesting VoxCPM's specific strength is cloning consistency, with IndexTTS2 holding an edge in Chinese prosodic naturalness specifically.

Getting Started with VoxCPM

  1. Check your environment requirements first. VoxCPM needs Python 3.10 or newer (below 3.13), PyTorch 2.5.0 or later, and CUDA 12.0 or newer — confirm these before attempting installation.
  2. Install the package and pull a checkpoint from Hugging Face. openbmb/VoxCPM-0.5B is the original bilingual release; openbmb/VoxCPM1.5 is the December 2025 update with SFT and LoRA fine-tuning support and improved 44.1kHz audio output, upgraded from the original's 16kHz.
  3. Load the model through the voxcpm Python package. from voxcpm import VoxCPM, then VoxCPM.from_pretrained("openbmb/VoxCPM-0.5B") gets you a ready-to-use model object in a couple of lines.
  4. Generate audio with model.generate(). Pass your target text along with generation parameters like cfg_value (classifier-free guidance strength) and inference_timesteps, then write the returned waveform out with a library like soundfile.
  5. Use ModelScope as an alternative download source if Hugging Face access is slow for you. The project documents this as a direct alternative for pulling the same weights.
  6. Clone the GitHub repository (OpenBMB/VoxCPM) for the full quick-start documentation, demo scripts, and licensing details before building anything beyond initial testing.

Tips for Better Results

  • Let context-aware generation do the work instead of over-specifying style. Since the model already infers prosody and emotion from text meaning, writing naturally descriptive input text tends to produce better results than trying to force a specific delivery through workarounds the model wasn't built to accept.
  • Use VoxCPM1.5 rather than the original 0.5B release for new projects. The upgraded 44.1kHz audio output and added fine-tuning support make it the more practical starting point unless you have a specific reason to use the earlier version.
  • Fine-tune with LoRA for a specific voice or domain rather than full retraining. Given the documented LoRA support, this is the more resource-efficient path if you need the model adapted to a particular voice or content style beyond zero-shot cloning.
  • Treat this as research-grade software, not a production-ready commercial product yet. The project's own model card explicitly states it's released for research and development purposes, and recommends against production or commercial use without rigorous independent testing and safety evaluation first.
  • Stick to Chinese and English for reliable results. Performance on other languages isn't guaranteed and can produce unpredictable or low-quality output, since the model's core training data is specifically bilingual rather than broadly multilingual.

VoxCPM vs. Other Strong Chinese-English Cloning Models

VoxCPMIndexTTS2CosyVoice2
Core representationTokenizer-free, continuous + FSQ semi-discreteDiscrete tokensSupervised semantic tokens
English WER (SEED-TTS-EVAL)1.85%HigherHigher
Chinese CER (SEED-TTS-EVAL)0.93%HigherHigher
Chinese naturalness (subjective)Strong, slightly behind IndexTTS2LeadingStrong
Speaker similarityLeading on English and ChineseCompetitiveCompetitive
Style/emotion controlAutomatic, context-inferredDuration control + emotion-timbre disentanglementInstruction-based
LicenseApache-2.0Bilibili model license (commercial use needs authorization)Open, hosted API available

VoxCPM's specific edge is a genuinely different architectural bet — continuous representation over discrete tokenization — that pays off directly in benchmark numbers and low-latency streaming, making it a strong technical choice specifically for bilingual Chinese-English work, even though its lack of direct manual style control is a real trade-off against models built around explicit instruction-based delivery.

Frequently Asked Questions

What does "tokenizer-free" mean for VoxCPM?

It means VoxCPM generates continuous speech representations directly, rather than first compressing speech into a discrete token vocabulary the way most comparable TTS systems do — an approach intended to avoid the information loss discrete tokenization can introduce.

Can I control emotion or speaking style directly in VoxCPM?

Not through explicit manual controls in the current version. The model infers appropriate prosody and emotion automatically from the meaning of the input text instead, which the project's own documentation notes as a current limitation for anyone who specifically needs direct style control.

Is VoxCPM ready for production or commercial use?

The project's own model card recommends against this without rigorous independent testing and safety evaluation first, explicitly describing the release as being for research and development purposes.

Does VoxCPM support languages beyond Chinese and English?

Not reliably. It's trained specifically as a bilingual Chinese-English model, and performance on other languages isn't guaranteed. A separate, newer release called VoxCPM2 expands to 2 billion parameters and 30 languages with additional voice design capabilities, if broader language coverage is what you need.

How fast is VoxCPM?

It achieves a real-time factor as low as 0.17 on a consumer-grade RTX 4090, meaning generation runs roughly six times faster than the resulting audio's own playback length.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →