EchoraEchora
Back to Models

FireRedTTS: A Foundation Model Built for Two Very Different Jobs

September 12, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Xiaohongshu's FireRed Team didn't set out to build a narrow demo model — they built a foundation framework, with an explicit data pipeline and two concrete downstream applications proving it out: cloning a voice for dubbing, and generating casual, human-like speech for a chatbot. Those are genuinely different problems (one needs to sound like a specific person, the other needs to sound like a natural conversation partner), and FireRedTTS was built specifically to show the same foundation system could do both well.

What Is FireRedTTS?

FireRedTTS is an open-source, language-model-based text-to-speech framework developed by Xiaohongshu's FireRed Team, introduced in a September 2024 paper. It's structured as three connected parts: a data processing pipeline that transforms raw audio into a large-scale, richly annotated TTS dataset; a foundation TTS system built on that data; and demonstrated downstream applications proving the foundation model's practical range. That three-part framing — data, foundation, applications — is a more deliberate, industry-oriented structure than a typical single-purpose research release.

Functional Advantages of FireRedTTS

  • A genuine foundation system, not a narrow single-task model. The same underlying architecture supports both zero-shot cloning and instruction-tuned conversational speech generation.
  • Solid in-context learning. The model stably synthesizes speech consistent with both the prompt text and the prompt audio's style, the same broad capability that makes large language models useful for tasks they weren't explicitly fine-tuned on.
  • Few-shot fine-tuning for studio-level quality, reaching expressive, professional-grade voice characters with just an hour of target-speaker recording.
  • Built transparently on established open components — including a text tokenizer approach from Fish Speech, the BigCodec speech codec, Encodec's causal convolution design, SpeechBrain's ECAPA-TDNN, and a Chinese-pretrained HuBERT model.

How the Foundation System Works

FireRedTTS compresses speech into discrete semantic tokens using a semantic-aware speech tokenizer, and a language model generates those tokens directly from prompt text and prompt audio — the in-context learning mechanism that lets the model pick up a target voice's style from the reference alone. A two-stage waveform generator then decodes those semantic tokens into a high-fidelity audio waveform. The codebase actually ships two different decoder configurations you can choose between: an acoustic LLM decoder, and a flow-matching decoder — giving you a choice between generation approaches depending on which trade-offs (speed, stability, quality character) fit your specific use case better.

Two Real-World Applications: Dubbing and Chatbots

This is where FireRedTTS's foundation-model framing actually pays off in practice. For dubbing, the model supports zero-shot voice cloning suited to UGC (user-generated content) scenarios — cloning a target voice on the fly with no additional training — and also supports few-shot fine-tuning with about an hour of target-speaker recording to reach studio-level expressive voice quality for PUGC (professional user-generated content), where higher production polish matters. For chatbots, FireRedTTS uses instruction tuning to produce controllable, human-like speech in a casual register, including paralinguistic behaviors and emotional coloring — tuned specifically for the register a spoken conversational agent needs, rather than the more formal delivery dubbing work typically calls for.

The FireRedTTS Family

The original 2024 release has continued to grow into a broader lineage. FireRedTTS-1S (2025) upgraded the system to genuinely streaming generation, offering a choice between a lower-latency multi-stream language-model approach and a lower-real-time-factor flow-matching approach depending on which metric matters more for your deployment. FireRedTTS-2 (2025) extended the architecture specifically toward long-form, multi-speaker dialogue generation for podcast and chatbot use cases, adding a 12.5Hz streaming tokenizer and reporting industry-leading results against comparable dialogue-synthesis systems. FireRedTTS3 (2026) pushed further into unified generation and editing, adding instruction-controlled voice design and the ability to precisely edit specific spans of existing speech. If your project specifically needs streaming, multi-speaker podcast-style dialogue, or speech editing, one of these later releases is likely the more directly relevant starting point than the original foundation model.

Getting Started with FireRedTTS

  1. Create a dedicated conda environment. conda create --name redtts python=3.10 sets up the required Python version before installing anything else.
  2. Install PyTorch matching your CUDA version. The project documents specific install commands for CUDA 11.8 and CUDA 12.1 — using the wrong one is a common source of runtime errors, so match this to your actual driver setup.
  3. Install FireRedTTS from source. cd fireredtts && pip install -e . followed by pip install -r requirements.txt completes the core installation.
  4. Download model files from the project's Model_Lists into pretrained_models.
  5. Choose your decoder in code. from fireredtts.fireredtts import FireRedTTS, then initialize with config_path="configs/config_24k.json" for the acoustic LLM decoder, or config_path="configs/config_24k_flow.json" for the flow-matching decoder — pick based on whether you're prioritizing generation stability or the specific quality characteristics of flow-matching output.

Tips for Better Results

  • Use zero-shot cloning for quick UGC-style dubbing, and few-shot fine-tuning for anything studio-facing. The roughly one-hour fine-tuning path is specifically documented for reaching expressive, professional-grade voice quality — don't expect zero-shot alone to match that polish for premium content.
  • Lean into instruction tuning for chatbot-style output. If your goal is a conversational agent's voice rather than narration or dubbing, the casual-register, paralinguistic-aware capability is what FireRedTTS was specifically tuned to deliver for that use case.
  • Pick your decoder deliberately, not by default. Test both the acoustic LLM decoder and the flow-matching decoder on your actual content, since the project ships both specifically because they represent different trade-offs rather than one being a strict upgrade.

FireRedTTS vs. Other Chinese-Team Foundation TTS Systems

FireRedTTSIndexTTSCosyVoice
Backing organizationXiaohongshu (FireRed Team)BilibiliAlibaba (FunAudioLLM)
Core architectureSemantic tokenizer + LM + two-stage waveform decoderXTTS/Tortoise-based, GPT-styleSupervised semantic tokens + flow matching
Demonstrated applicationsDubbing (zero-shot + few-shot) and chatbot speechGeneral zero-shot cloning, Chinese pronunciation accuracyZero-shot multilingual cloning
Fine-tuning for studio qualityYes, ~1 hour of dataNot the primary focusNot the primary focus
License / use restrictionsAcademic research use stated explicitlyWeights need authorization for commercial useOpen, hosted API available

FireRedTTS's specific niche is being explicitly validated across two genuinely different downstream jobs from one foundation system — most comparable Chinese-team TTS releases lead with a single primary use case, while FireRedTTS's own paper frames dubbing and chatbot speech as equally central proof points.

Frequently Asked Questions

What's the difference between FireRedTTS and later releases like FireRedTTS-2?

The original FireRedTTS is the foundation system, demonstrated for dubbing and chatbot speech. FireRedTTS-2 specifically targets long-form, multi-speaker dialogue generation for podcasts, and FireRedTTS3 adds unified generation and speech editing — check which release matches your actual use case before starting with the original.

Can FireRedTTS clone a voice with no additional training?

Yes, through zero-shot cloning suited to quick, UGC-style dubbing scenarios. For higher, studio-level expressive quality, the project also documents a few-shot fine-tuning path using about an hour of target-speaker recording.

Is FireRedTTS free for commercial use?

Its zero-shot voice cloning functionality is intended solely for academic research purposes and explicitly prohibits illegal use — review current license and usage terms directly before any commercial deployment.

Which decoder should I use, the acoustic LLM decoder or the flow-matching decoder?

Both are provided because they represent different trade-offs rather than one being strictly better — test both on your specific content to see which fits your priorities around generation stability and output quality character.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →