EchoraEchora
Back to Models

F5-TTS: Text In, Speech Out, No Extra Machinery Required

August 29, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Most voice-cloning pipelines carry more moving parts than they need to: a duration predictor to guess how long the output should be, a phoneme aligner to map text to sound, a dedicated text encoder trained separately from the rest of the system. F5-TTS strips all of that out. It pads the input text to match the target speech length and lets a Diffusion Transformer handle the rest — a genuinely simpler pipeline that also happens to be fast, with a real-time inference factor of just 0.15 in the researchers' own testing.

What Is F5-TTS?

F5-TTS is a fully non-autoregressive text-to-speech system built on flow matching with a Diffusion Transformer (DiT) backbone, developed by researchers from Shanghai Jiao Tong University, the University of Cambridge, and Geely Automobile Research Institute, and released as open source under the name SWivid/F5-TTS. Rather than predicting speech token by token the way autoregressive models do, F5-TTS treats generation as a denoising process — starting from noise and using the DiT to iteratively transform it into a mel spectrogram conditioned on the input text and a reference voice.

Under the Hood: Why F5-TTS Is Fast

Filler-token padding instead of duration modeling. Rather than predicting how long each word should take, F5-TTS simply pads the text input with filler tokens until it matches the length of the target speech, then lets the flow-matching process figure out the alignment during denoising. This eliminates an entire component — the duration model — that most TTS architectures treat as essential.

ConvNeXt V2 for text representation. This is one of F5-TTS's key refinements over its direct predecessor. Earlier flow-matching approaches struggled with reliably aligning text and speech, leading to skipped words or repeated phrases. F5-TTS runs the text through ConvNeXt V2 first, producing a representation that's meaningfully easier for the diffusion transformer to align with the corresponding audio.

Sway Sampling. An inference-time strategy that adjusts how flow steps are sampled during generation, improving both output quality and inference speed without any retraining. Notably, the technique isn't proprietary to F5-TTS's specific architecture — it generalizes to other conditional flow-matching models as a drop-in inference improvement.

The combined effect of these choices is a genuinely lean pipeline: no duration model, no phoneme aligner, no separately trained text encoder, and a reported real-time factor of 0.15 — meaning generation runs roughly 6-7x faster than the audio's own playback length, competitive with heavily optimized production TTS systems.

The E2 TTS Connection

F5-TTS didn't invent the "pad text, then denoise" approach from scratch — it directly builds on E2 TTS, an earlier model that first proved this minimal, aligner-free strategy could work at all. E2 TTS demonstrated the concept was viable but suffered from slow convergence during training and alignment robustness issues that made it difficult to deploy reliably. F5-TTS was built specifically to fix those problems, using ConvNeXt V2 and Sway Sampling to solve exactly the failure modes E2 TTS exposed. The official F5-TTS repository actually ships an E2TTS_Base checkpoint alongside the F5 models specifically so researchers can reproduce and compare both architectures directly, rather than relying on the original paper's reported numbers alone.

Zero-Shot and Cross-Lingual Voice Cloning

F5-TTS voice cloning works from a short reference clip — commonly cited in the 5 to 15 second range, with some setups working from as little as 6 seconds — with no fine-tuning step required to get a usable result. It also supports cross-lingual cloning: provide an English reference clip and generate Spanish (or another supported language) output in that same cloned voice, a capability the model's own documentation highlights specifically.

For teams that need more consistency than zero-shot delivers, F5-TTS also supports fine-tuning, and its training loop is comparatively simple — no codec preprocessing step and no multi-stage pipeline to manage, just prepared audio and aligned transcripts. That relative simplicity is part of why it's a common starting point for teams experimenting with custom voice training on a single consumer GPU rather than a full training cluster.

Getting Started with F5-TTS

  1. Clone the repository. git clone https://github.com/SWivid/F5-TTS.git, then cd F5-TTS and pip install -e . — add the optional recursive submodule step only if you plan to use BigVGAN as your vocoder.
  2. Install PyTorch matching your CUDA version. The project's documentation specifies exact PyTorch and torchaudio versions per CUDA release; matching these precisely avoids most installation issues.
  3. Run inference via CLI. f5-tts_infer-cli --model F5TTS_v1_Base --ref_audio "your_clip.wav" --ref_text "transcript of the clip" --gen_text "the text you want spoken" generates a clip directly; leaving --ref_text blank lets an ASR model transcribe the reference automatically, at the cost of extra GPU memory.
  4. Use the Gradio WebUI for quick testing. f5-tts_infer-gradio (or the equivalent Docker command) opens a browser interface for testing cloning before wiring anything into a pipeline.
  5. Check your VRAM against recommended specs. F5-TTS runs on any NVIDIA GPU with roughly 16GB of VRAM, with 24GB (RTX 3090/4090-class cards) providing comfortable headroom — confirm your hardware fits before planning a production deployment.
  6. Use the TensorRT-LLM/Triton path for production-scale serving. The project documents a deployment route using Triton and TensorRT-LLM for significantly faster decoding when serving many requests, rather than running the base Python inference script at scale.

Tips for Better Results

  • Keep your reference clip clean and in the recommended length range. Background noise or unusually short/long clips outside the 5–15 second window tend to degrade cloning quality more than they would with a purely acoustic-token model.
  • Provide an accurate reference transcript when possible. Letting the built-in ASR model auto-transcribe is convenient, but a manually verified transcript avoids compounding any ASR errors into the cloning process itself.
  • Test cross-lingual cloning on your specific language pair before committing. The capability is real and documented, but quality still varies by the specific source-to-target language combination, so validate the pair your project actually needs.
  • Use Sway Sampling's default settings before tuning anything manually. It's designed to improve results out of the box without retraining — treat manual flow-step tuning as a later optimization step, not a starting requirement.
  • Read the license terms before any commercial use. F5-TTS's pretrained weights are released under a CC-BY-NC license, a direct consequence of being trained on the Emilia in-the-wild dataset — this permits research and non-commercial use, but commercial deployment needs separate clearance.

F5-TTS vs. E2 TTS vs. Other Cloning Models

F5-TTSE2 TTS (predecessor)CosyVoice3
ArchitectureDiffusion Transformer + flow matchingFlat U-Net TransformerSupervised semantic tokens + flow matching
Duration model / aligner neededNoNo (but less robust)No
Training convergenceFaster, more robustSlow, alignment issuesN/A (different architecture family)
Reference clip length~5–15 secondsSimilar~3–10 seconds
Cross-lingual cloningYesLimitedYes
Inference speed (RTF)~0.15SlowerComparable, hosted variants ~150ms latency
LicenseCC-BY-NC (non-commercial)Research useOpen, hosted API available

F5-TTS's specific contribution to this family is proving that the minimal, aligner-free approach E2 TTS pioneered could actually be made fast and reliable — the architecture is simpler than most competing cloning models, and the benchmark speed reflects that simplicity directly rather than coming from extra optimization layered on top.

Frequently Asked Questions

What does "non-autoregressive" mean for F5-TTS, practically?

It means the model doesn't generate speech one token at a time in sequence. Instead, it starts from noise and denoises the entire output through the Diffusion Transformer, which is part of why inference can be faster than token-by-token autoregressive generation.

What's the relationship between F5-TTS and E2 TTS?

E2 TTS was the earlier model that first proved a duration-model-free, aligner-free approach to TTS could work, but it suffered from slow training convergence and alignment robustness issues. F5-TTS was built specifically to fix those problems using ConvNeXt V2 text representation and Sway Sampling, and the official repository includes both F5-TTS and E2TTS_Base checkpoints for direct comparison.

How much reference audio does F5-TTS need for voice cloning?

Commonly cited guidance is 5 to 15 seconds, though some setups report usable results from as little as 6 seconds of clean reference audio.

Is F5-TTS free for commercial use?

Not without separate clearance. The pretrained model weights are released under a CC-BY-NC license because of the Emilia training dataset's licensing terms — this covers research and non-commercial use, but commercial deployment requires checking current terms directly.

What hardware do I need to run F5-TTS?

Any NVIDIA GPU with roughly 16GB of VRAM works, with 24GB (RTX 3090 or 4090-class cards) recommended for comfortable headroom, particularly if you plan to fine-tune.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →