IndexTTS: The Voice Model Built Inside China's Largest Anime and Gaming Video Platform
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Bilibili — the video platform behind millions of hours of anime, gaming, and creator content in China — has an obvious, practical reason to care about text-to-speech: dubbing, digital avatars, and voiceover at platform scale. IndexTTS is the result of that internal need going public. Built by Bilibili's AI Platform Department and released as open source, it's a zero-shot voice cloning model engineered specifically around a problem that generic multilingual TTS systems tend to handle poorly — correctly pronouncing Chinese text with all its polyphonic characters and long-tail vocabulary.
What Is IndexTTS?
IndexTTS is a GPT-style, autoregressive text-to-speech model developed by Bilibili's Index team, built primarily on the XTTS and Tortoise architectures with a set of targeted improvements layered on top. It handles zero-shot voice cloning — reproducing a speaker's voice from a reference clip — in both Chinese and English, and was trained on more than ten thousand hours of speech data.
What distinguishes IndexTTS from the XTTS and Tortoise foundations it builds on isn't a single headline feature so much as a series of deliberate engineering choices, each aimed at a specific weak point in earlier open-source TTS: pronunciation accuracy for Chinese, controllable timing, and cloning stability.
The Engineering Behind IndexTTS
Character-pinyin hybrid modeling. This is IndexTTS's signature contribution. Chinese text synthesis has a persistent problem: polyphonic characters (the same character read differently depending on context) and long-tail vocabulary frequently get mispronounced by models trained purely on character sequences. IndexTTS addresses this by mixing characters and pinyin directly in training — during preprocessing, a portion of non-polyphonic characters are randomly swapped for their pinyin equivalents, teaching the model to use pinyin as a disambiguation signal it can also accept at inference time to force a specific pronunciation.
Punctuation-based pause control. IndexTTS reads punctuation marks as timing signals, letting you control where pauses occur in generated speech simply through how you punctuate the input text — no separate timing markup required.
A rebuilt conditioning and decoding pipeline. IndexTTS introduces a conformer-based speech conditional encoder to improve how the model captures speaker characteristics from a reference clip, and replaces the original speech-code decoder with BigVGAN2, a change the team credits with improved training stability, timbre similarity, and overall audio quality.
FSQ over VQ for the speech tokenizer. The team ran a direct comparison between standard Vector Quantization and Finite-Scalar Quantization for encoding acoustic speech tokens, addressing the codebook collapse issue that can degrade VQ-based systems — a technical choice that later influenced the tokenizer design in CosyVoice2 as well.
Benchmark Performance
In the team's published evaluation — measuring word error rate, speaker similarity, and mean opinion score across English and Chinese test sets — IndexTTS was compared directly against XTTS, CosyVoice2 (non-streaming), Fish-Speech, FireRedTTS, and F5-TTS. The reported results show IndexTTS outperforming all of them on word error rate, while remaining competitive with CosyVoice2 and F5-TTS specifically on speaker similarity — the two models the authors identify as its closest peers on cloning fidelity.
Worth noting honestly: the original IndexTTS paper acknowledges its own limitations directly — at release, the model didn't support instructed voice generation, covered only Chinese and English, and had limited capacity for rich emotional expression. The team's IndexTTS-1.5 update, released in May 2025, specifically targeted stability and English-language performance improvements over the initial 1.0 release.
Getting Started with IndexTTS
- Set up a Python 3.10 environment. Create a dedicated conda environment and install
ffmpeg, either through your system package manager or viaconda install -c conda-forge ffmpeg. - Handle the Windows
pyninidependency separately if needed. Ifpip installfails building a wheel forpyninion Windows, install it viaconda install -c conda-forge pynini==2.1.6followed bypip install WeTextProcessing --no-deps— a known workaround the project documents directly. - Download model weights. Use
huggingface-cli downloadto pull eitherIndexTeam/IndexTTS(1.0) orIndexTeam/IndexTTS-1.5, depending on which release you need — the 1.5 checkpoint is the better default for most new projects given its stability improvements. - Run inference via CLI or Python. The
indexttscommand-line tool takes text and a reference voice file directly, or you can call theIndexTTSclass from Python for more control within a larger pipeline. - Use the WebUI to test before integrating.
pip install -e ".[webui]" --no-build-isolationfollowed bypython webui.pyopens a browser-based interface atlocalhost:7860for testing cloning and pronunciation control without writing code.
Tips for Better Results
- Use pinyin overrides deliberately for polyphonic characters. When a character has multiple valid pronunciations depending on context, supplying the correct pinyin directly is more reliable than hoping the model infers the right reading from surrounding text.
- Punctuate for the pauses you actually want. Since IndexTTS reads punctuation as timing information, treat comma and period placement as a pacing tool, not just grammatical correctness — this is a lower-effort way to control rhythm than most models offer.
- Default to IndexTTS-1.5 unless you have a specific reason not to. The 1.5 release exists specifically to fix stability issues and improve English output from the original 1.0 model — there's little reason to start a new project on the earlier checkpoint.
- Keep expectations realistic on emotional range. The model's own documentation acknowledges limited emotional expressiveness compared to more recent, emotion-focused releases — if dramatic delivery is central to your project, this generation of IndexTTS isn't optimized for it.
- Check the model license before commercial deployment. IndexTTS's code is released under Apache 2.0, but the pretrained model weights carry a separate Bilibili model license that requires prior written authorization for commercial use — confirm your use case is covered before shipping a commercial product on it.
IndexTTS vs. Its Closest Peers
| IndexTTS | XTTS | CosyVoice2 | F5-TTS | |
|---|---|---|---|---|
| Architecture basis | XTTS + Tortoise, GPT-style | Baseline | Supervised semantic tokens | Flow matching |
| Chinese pronunciation control | Character-pinyin hybrid modeling | Not a focus | Not a dedicated feature | Not a dedicated feature |
| Pause control | Via punctuation | Limited | Via instruction | Limited |
| Word error rate (published eval) | Best among compared models | Baseline | Close second | Close second |
| Speaker similarity | Competitive with top models | Lower | Comparable | Comparable |
| Commercial use | Requires authorization for weights | Varies | Apache 2.0 | Varies by version |
IndexTTS's specific edge is Chinese pronunciation accuracy and inference simplicity relative to its closest competitors — the team describes its training process as comparatively straightforward and its usage as more controllable than several open-source peers, which is a meaningful practical advantage for teams building production pipelines rather than research demos.
Frequently Asked Questions
Who developed IndexTTS?
IndexTTS was developed by the AI Platform Department at Bilibili (B站), China's major anime, gaming, and creator video platform, and released as open source on GitHub and Hugging Face.
What's the difference between IndexTTS-1.0 and IndexTTS-1.5?
IndexTTS-1.5, released in May 2025, is an update focused on improving model stability and English-language synthesis quality over the original 1.0 release — most new projects should default to 1.5.
How does IndexTTS handle Chinese pronunciation issues?
Through character-pinyin hybrid modeling — the model is trained on a mix of Chinese characters and their pinyin equivalents, which lets you supply the correct pinyin directly to fix a mispronounced polyphonic or uncommon character.
Is IndexTTS free for commercial use?
The code is released under the Apache 2.0 license, but the pretrained model weights carry a separate Bilibili license that requires prior written authorization for commercial use — check the current license terms before deploying commercially.
How does IndexTTS compare to CosyVoice2 and F5-TTS?
In the team's published benchmarks, IndexTTS achieved a lower word error rate than both, while remaining competitive with them specifically on speaker similarity — the metric that measures how closely cloned output matches the reference voice.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.