EchoraEchora
Back to Models

CosyVoice3: The Version Built for Every Accent in the Room

August 28, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

CosyVoice2 solved latency — streaming synthesis at 150ms first-packet delay made it usable for real-time voice chat. What it didn't fully solve was breadth: a voice product serving Chinese-speaking users doesn't just need Mandarin, it needs Cantonese, Sichuanese, Shanghainese, and a dozen other regional accents to sound authentic rather than generic. CosyVoice3 is Alibaba FunAudioLLM's answer to that gap — the same streaming foundation, extended to cover more languages, far more Chinese dialects, and a new way to fix pronunciation without retraining anything.

What's New in CosyVoice3

Built around the same 0.5B-parameter scale as CosyVoice2, CosyVoice3 (released as Fun-CosyVoice3-0.5B) is a refinement release rather than a ground-up rebuild — it targets three specific weaknesses the team identified in the previous generation: content consistency, speaker similarity, and prosody naturalness. The result, according to Alibaba's published benchmarks, is a model that outperforms several larger competing systems despite staying at the same compact parameter count.

Dialect and Language Coverage, Substantially Expanded

This is the headline change, and it's the reason CosyVoice3 exists as a distinct release rather than a minor version bump. CosyVoice3 supports nine languages — Chinese, English, Japanese, Korean, German, Spanish, French, Italian, and Russian — alongside more than eighteen Chinese regional dialects and accents, including Cantonese, Sichuanese, Shanghainese, Northeastern Mandarin, Shaanxi, Shanxi, Tianjin, Shandong, Ningxia, and Gansu varieties. Few open-source TTS models attempt this level of dialect granularity within one checkpoint, and it's what makes CosyVoice3 specifically well-suited to Chinese-market products where a single "Mandarin" voice isn't a convincing substitute for a listener's actual regional accent.

Pronunciation Inpainting

CosyVoice3 introduces a capability the earlier versions didn't have: the ability to correct a specific mispronounced word — commonly a polyphonic Chinese character or an uncommon term — by supplying the correct Pinyin or phoneme sequence directly, without retraining or fine-tuning the model. For production use in domains like medical, legal, or technical content, where one wrong reading can undermine an entire clip, this turns what used to be a dead-end limitation into a fixable parameter.

Improved Cross-Lingual Cloning

Cloning a voice from a Chinese reference clip and generating natural English (or the reverse) was already possible in earlier versions, but CosyVoice3 specifically targeted this scenario for improvement — the team's stated goal was tighter voice-identity preservation when the reference language and generation language don't match.

The Benchmark Numbers

Alibaba's published evaluations position CosyVoice3 favorably against both its own predecessor and larger third-party models:

  • Chinese character error rate of 0.81% on its RL-tuned variant — lower than F5-TTS's 1.52% and VibeVoice-1.5B's 1.16%, despite CosyVoice3 running at roughly a third of VibeVoice's parameter count.
  • English word error rate of 1.68%, again ahead of F5-TTS (2.00%) and VibeVoice-1.5B (3.04%).
  • Speaker similarity of 78.0% in Chinese and 71.8% in English, with the Chinese figure reported as exceeding the similarity benchmark set by actual human re-recordings in the same evaluation.

The practical takeaway: parameter count isn't the deciding factor here. CosyVoice3's advantage comes from the specific architectural refinements — better token representation, improved training, and targeted dialect data — rather than simply scaling up.

Getting Started with CosyVoice3

  1. Clone the repository with submodules. CosyVoice3 shares the same GitHub repository as earlier versions: git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git — the --recursive flag is required for the Matcha-TTS dependency.
  2. Download Fun-CosyVoice3-0.5B specifically. Use ModelScope's snapshot_download or pull from Hugging Face — make sure you're targeting the 3.0 checkpoint rather than an older CosyVoice2 or original CosyVoice model directory.
  3. Select a dialect or language explicitly when it matters. If your use case needs a specific regional accent, choose the corresponding dialect voice or instruction rather than relying on the general Chinese model to infer it from context.
  4. Launch the WebUI to test before integrating. python3 webui.py --port 50000 --model_dir pretrained_models/Fun-CosyVoice3-0.5B opens a Gradio interface for testing cloning, dialect selection, and pronunciation inpainting directly.
  5. Use TensorRT-LLM if inference speed becomes a bottleneck at scale. The project reports roughly 4x acceleration for the language-model component compared to standard Transformers inference — worth the setup effort for production deployments handling high request volume.
  6. Consider the hosted API for teams that don't want to self-host. Alibaba Cloud's Model Studio offers CosyVoice3 as a managed API (cosyvoice-v3-flash / cosyvoice-v3-plus) with quoted first-package latency as low as 150ms, if running your own inference stack isn't a priority.

Tips for Better Results

  • Reach for pronunciation inpainting before assuming a word is unfixable. A mispronounced polyphonic character or technical term is usually a one-line phoneme correction, not a reason to regenerate the whole clip or abandon the model for that content.
  • Be explicit about dialect rather than implicit. With 18+ dialects available, letting the model guess from context is less reliable than specifying the target accent directly, especially for less common dialects like Ningxia or Gansu.
  • Validate cross-lingual pairs individually. The cross-lingual cloning improvement is real, but quality still varies by specific language pair — test the exact source-to-target combination your product needs rather than assuming uniform results across all nine languages.
  • Use clean 3–10 second reference clips. This remains the window the model's zero-shot cloning is tuned around; noisy or unusually long reference audio degrades speaker similarity regardless of which version you're running.
  • Don't default to the largest available model out of habit. CosyVoice3's benchmark results outperforming larger models like VibeVoice-1.5B are a reminder that architecture and training data matter more than raw parameter count — evaluate on your actual content before assuming a bigger model elsewhere would do better.

CosyVoice3 vs. CosyVoice2 vs. Original CosyVoice

CosyVoice3CosyVoice2CosyVoice (original)
Parameters0.5B0.5B300M
Languages9Multilingual (fewer dialects)5
Chinese dialects18+LimitedNot a focus
Pronunciation inpaintingYesNoNo
StreamingYesYes, introduced hereNo
Cross-lingual cloningImprovedSupportedSupported
Chinese CER (best variant)0.81%HigherBaseline

Each version solved a different constraint: the original proved supervised tokens beat unsupervised ones, CosyVoice2 added the streaming architecture needed for real-time use, and CosyVoice3 extends that same foundation across far more languages and dialects while sharpening pronunciation accuracy. If your product specifically needs authentic regional Chinese accents or a way to fix mispronunciations without retraining, CosyVoice3 is the version built for that.

Frequently Asked Questions

What's the main upgrade in CosyVoice3 over CosyVoice2?

Primarily dialect and language breadth — CosyVoice3 covers nine languages and more than eighteen Chinese regional dialects, versus more limited coverage in CosyVoice2 — plus the new pronunciation inpainting feature and improved cross-lingual cloning accuracy.

What is pronunciation inpainting?

A CosyVoice3 feature that lets you correct a specific mispronounced word — typically a polyphonic Chinese character or uncommon term — by supplying the correct Pinyin or phoneme sequence directly, without retraining the model.

Does CosyVoice3 still support streaming, low-latency generation?

Yes. CosyVoice3 builds on the streaming architecture introduced in CosyVoice2, with hosted API options quoting first-package latency as low as 150ms.

How does CosyVoice3 compare to larger models like VibeVoice or F5-TTS?

In Alibaba's published benchmarks, CosyVoice3 outperformed both on Chinese and English error rates and voice-cloning speaker similarity, despite running at a smaller parameter count than VibeVoice-1.5B specifically.

Is CosyVoice3 free and open-source?

Yes. Model weights and code are available through the FunAudioLLM/CosyVoice GitHub repository and ModelScope, alongside a separate hosted API through Alibaba Cloud for teams that prefer a managed service.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →