IndexTTS2: The First TTS Model That Can Match a Lip Flap
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Dub a video with most TTS models and you'll hit the same wall eventually: the generated line is half a second too long, and now the audio drifts out of sync with the character's mouth on screen. Autoregressive TTS models — the kind that generate speech token by token and tend to sound the most natural — have historically been the worst offenders here, because their generation length isn't something you can pin down in advance. IndexTTS2, from Bilibili's Index team, is the first autoregressive TTS model built to solve that specific problem: precise, millisecond-level control over how long a line of generated speech actually takes.
What's New in IndexTTS2
IndexTTS 2 is Bilibili's follow-up to the original IndexTTS, and it's less an incremental update than a rebuild aimed squarely at professional dubbing and content-localization workflows. Where the original model explicitly acknowledged limited emotional range and no instructed generation as open problems, IndexTTS2 addresses both directly — while adding a capability the original never had at all: control over exactly how long a generated line lasts.
Trained on roughly 55,000 hours of multilingual speech data spanning Chinese, English, and Japanese, IndexTTS2 is described by its authors as achieving state-of-the-art results against existing zero-shot TTS systems on word error rate, speaker similarity, and emotional fidelity simultaneously — three metrics that are often improved at each other's expense in other models.
Precise Duration Control: The Headline Feature
This is IndexTTS2's core contribution, and it directly targets a structural limitation of autoregressive TTS. Because these models generate speech one token at a time without a fixed target length, matching a specific duration — say, exactly 3.2 seconds to fit a dubbed line into an existing video clip — has traditionally required trial and error, re-generation, or post-processing time-stretching that degrades audio quality.
IndexTTS2 introduces a duration adaptation scheme with two distinct generation modes:
- Specified duration mode. Explicitly set the number of generated tokens to hit a precise target length — the mode built for audiovisual dubbing, where a line has to land within an exact time window to match lip movement or scene timing.
- Free duration mode. Let the model generate naturally without a token-count constraint, while still faithfully reproducing the prosodic characteristics of the input prompt — the better choice when natural pacing matters more than hitting an exact duration.
This is, notably, a general method the authors designed to be applicable to other autoregressive TTS architectures beyond IndexTTS2 itself — not a narrow trick specific to this one model.
Emotion-Timbre Disentanglement
The second major addition is decoupling who a voice sounds like from how it feels. In most cloning models, the reference clip you provide determines both the speaker's identity and, implicitly, their emotional tone — you can't easily clone a calm speaker's voice and have it deliver an angry line convincingly. IndexTTS2 separates these two dimensions, letting you specify a timbre source and an emotional source independently, including from two entirely different reference clips.
Emotional input is supported through four separate methods: reference audio alone, a separate emotion-audio prompt paired with different target text, a plain-language emotional description (e.g., "sound anxious and rushed"), or a direct emotion vector. The natural-language description path is powered by a fine-tuned Qwen3 model acting as a soft instruction layer, specifically designed to lower the barrier for emotional control without requiring users to understand vector-based parameters. To keep speech clear during intense emotional delivery — a common failure point where strong emotion degrades intelligibility — the model incorporates GPT latent representations aimed at preserving stability under emotional stress.
Getting Started with IndexTTS2
- Clone the repository.
git clone https://github.com/index-tts/index-tts.git, ensuring bothgitandgit-lfsare installed on your system before proceeding. - Install dependencies with
uv. The project recommends theuvpackage manager over a manual pip/conda setup:uv sync --all-extraspulls in the full dependency set. - Download the IndexTTS2 checkpoint.
hf download IndexTeam/IndexTTS-2 --local-dir=checkpointsretrieves the model weights specifically — don't reuse an IndexTTS-1.5 checkpoint directory, as the two aren't interchangeable. - Use the v2 inference class. In Python, import from
indextts.infer_v2rather than the originalinfermodule, then call.infer()with a speaker audio prompt and target text to generate a basic clip. - Layer in duration or emotion control for dubbing work. Once basic generation is working, add a target duration or an emotion-audio prompt to your inference call — this is where IndexTTS2's dubbing-specific advantages actually show up, not in default single-shot generation.
Tips for Getting the Most Out of IndexTTS2
- Use specified duration mode for anything with a visual timing constraint. Free duration mode is fine for standalone narration, but any project syncing audio to existing video — dubbing, redubs, localized ads — should use the explicit token-count mode from the start rather than fixing timing in post.
- Separate your timbre and emotion sources deliberately. If you need a specific person's cloned voice to deliver a line with an emotional tone that reference clip doesn't naturally have, provide a separate emotion-audio prompt rather than trying to coax emotion out of a single flat reference clip.
- Reach for the text-description emotion input first if you're not an audio engineer. The Qwen3-powered natural-language instruction path is built to be more accessible than emotion vectors — start there before assuming you need to hand-tune numerical parameters.
- Watch clarity during high-emotion generation. While IndexTTS2's GPT latent representations are designed to preserve intelligibility under strong emotional delivery, it's still worth spot-checking clarity on your specific content rather than assuming it's uniform across every emotion and language combination.
- Confirm your license terms before commercial dubbing work. As with the original IndexTTS, model weights are governed by Bilibili's model use license, which requires authorization for commercial use — verify this covers your project before shipping a paid localization or dubbing product built on it.
IndexTTS2 vs. IndexTTS (Original)
| IndexTTS2 | IndexTTS (1.0 / 1.5) | |
|---|---|---|
| Duration control | Precise, millisecond-level (two modes) | Not supported |
| Emotion control | Disentangled from timbre, 4 input methods | Limited, acknowledged weakness |
| Instructed generation | Yes, via natural-language description | Not supported |
| Training data | ~55,000 hours, Chinese/English/Japanese | ~10,000+ hours, Chinese/English |
| Best fit | Video dubbing, localization, emotional narration | General zero-shot cloning, Chinese pronunciation accuracy |
| Released | September 2025 | March 2025 (1.0) / May 2025 (1.5) |
If your project is general-purpose voice cloning without strict timing requirements, the original IndexTTS remains a capable, simpler option. If you're dubbing video content where a line has to land within a specific time window — or you need a cloned voice to convincingly deliver an emotion the reference clip didn't have — IndexTTS2 is the version built specifically for that job.
Frequently Asked Questions
What does "duration control" mean in IndexTTS2, practically?
It means you can specify exactly how many tokens (and therefore roughly how much time) a generated line should take, which is essential for dubbing work where dialogue has to match existing video timing rather than running however long the model happens to generate.
Can IndexTTS2 make a cloned voice sound like a different emotion than the reference clip?
Yes — this is the emotion-timbre disentanglement feature. You can supply a separate emotion source (audio, text description, or emotion vector) independent of the timbre reference, so a calm reference voice can still deliver an anxious or excited line.
Is IndexTTS2 the same as "IndexTTS 2"?
Yes, they refer to the same model — IndexTTS2 (sometimes written IndexTTS 2) is Bilibili's second-generation release, distinct from the original IndexTTS (1.0/1.5).
What languages does IndexTTS2 support?
Its training data spans Chinese, English, and Japanese, drawn from roughly 55,000 hours of multilingual speech.
Is IndexTTS2 free for commercial use?
The code is open-sourced, but the model weights are governed by Bilibili's model use license agreement, which requires prior authorization for commercial use — check current license terms before deploying it in a paid product.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.