Tortoise TTS: The Slow Model That Everything Else Learned From
- What Is Tortoise TTS?
- The Architecture: An Autoregressive Transformer, Then a Diffusion Model
- Multi-Voice Cloning, the Deliberate Way
- Tortoise's Influence on Later Models
- Getting Started with Tortoise TTS
- Tips for Better Results
- Tortoise TTS vs. Newer Cloning Models
- Frequently Asked Questions
- Create Speech Online with Echora
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Before zero-shot cloning became fast and disposable, there was a model that made a different bet entirely: take as long as it needs, but make the output genuinely convincing. Tortoise TTS earned its name honestly — the developer picked it specifically because the model is, in his own words, insanely slow. What that patience buys you is speech quality that, back in 2022, was one of the first open-source demonstrations that neural TTS could go toe-to-toe with commercial systems, and its architectural DNA is still visible in newer models built partly on the ideas it proved out.
What Is Tortoise TTS?
Tortoise TTS is an open-source, multi-voice text-to-speech system created by developer James Betker, with its first commit dating back to January 2022. It's documented in Betker's paper "Better Speech Synthesis Through Scaling," which describes borrowing techniques that had recently transformed image generation — autoregressive transformers and denoising diffusion probabilistic models (DDPMs) — and applying them to speech instead.
Notably, Tortoise was built entirely by Betker on his own hardware, independent of any employer, and released under the fully permissive Apache 2.0 license — a genuinely commercial-friendly choice that stands out among several newer cloning models that restrict commercial use through non-commercial research licenses.
The Architecture: An Autoregressive Transformer, Then a Diffusion Model
Tortoise's approach is DALL-E-inspired, and the two-stage pipeline is the reason for both its strengths and its slowness. First, an autoregressive transformer — architecturally similar to GPT — generates discrete speech tokens conditioned on the input text and a reference voice sample. Rather than stopping there, Tortoise then hands those tokens to a diffusion model, which iteratively refines them into the final, high-fidelity audio waveform.
This is the same two-step logic that drove the leap in image generation quality a few years earlier: an initial generative pass followed by a refinement process that trades compute time for fidelity. In Tortoise's case, that trade-off produces genuinely excellent prosody and intonation — natural-sounding pauses, emphasis, and speaker-specific vocal texture — at the cost of generation speed that's noticeably slower than most modern alternatives.
Multi-Voice Cloning, the Deliberate Way
Tortoise supports voice cloning from a set of reference audio clips, and it's built specifically around handling multiple voices well rather than being optimized for a single default speaker. You provide several reference clips of a target voice, and the model uses them together to capture that speaker's characteristics — a slightly heavier setup than single-clip zero-shot cloning in newer models, but one that historically produced more stable, more natural-sounding identity transfer than the faster systems that followed it.
Tortoise's Influence on Later Models
This is worth knowing if you're evaluating Tortoise against newer options: its architectural ideas didn't stay contained to one project. XTTS's GPT-style autoregressive design and Bilibili's IndexTTS both explicitly build on foundations Tortoise helped establish — the combination of an autoregressive language model for content generation with a separate refinement stage for audio quality became a template that a meaningful part of the open-source TTS ecosystem adapted and iterated on. If you've used a newer cloning model that pairs a transformer with a separate acoustic refinement step, there's a reasonable chance some of that design lineage traces back here.
Getting Started with Tortoise TTS
- Set up a dedicated conda environment.
conda create --name tortoise python=3.9 numba inflect, then activate it and install PyTorch, torchvision, and torchaudio matched to your CUDA version. - Pin the transformers version exactly. Tortoise requires
transformers==4.29.2specifically — a newer version will likely break compatibility, so don't skip this pin during setup. - Clone the repository and install.
git clone https://github.com/neonbjb/tortoise-tts.git, thencd tortoise-ttsandpython setup.py install(or install the latest development version directly viapip install git+https://github.com/neonbjb/tortoise-tts). - Choose a speed/quality preset. Generation supports presets like
fast,standard, and higher-quality settings — start withfastfor initial testing before committing to slower, higher-fidelity presets for final output. - Enable
kv_cacheand half-precision for meaningfully faster generation. InitializingTextToSpeech(kv_cache=True)orTextToSpeech(half=True)speeds up inference noticeably compared to default settings, without a dramatic quality trade-off. - Try the hosted demo before committing to local setup. A Hugging Face Spaces demo is available — note that it requires a GPU-backed Space, since CPU-only instances won't run Tortoise's inference pipeline.
Tips for Better Results
- Budget real time for generation, especially at higher quality settings. Tortoise's name isn't ironic — even with
kv_cacheand half-precision enabled, this is not the model to reach for if you need near-instant turnaround; it's built for quality-first, offline workflows like audiobook narration rather than live interaction. - Provide several reference clips, not just one. Tortoise's multi-voice cloning was built around using multiple samples of a target voice together — leaning on a single short clip, the way newer zero-shot models are designed for, tends to undersell what it's actually capable of.
- Reserve it for content where quality matters more than turnaround. Long-form narration, audiobooks, and premium content are where Tortoise's prosody advantage is most noticeable — for live agents or rapid iteration, a newer low-latency model will serve you better.
- Confirm roughly 8GB of VRAM before planning a deployment. That's the commonly cited baseline for running Tortoise comfortably; check your specific preset and batch settings against that before assuming a smaller GPU will do.
- Don't expect the same install experience as an actively-developed model. Because Tortoise's development pace has slowed since its initial releases, sticking closely to documented dependency versions (especially that pinned
transformersrelease) will save you more troubleshooting time than assuming the latest versions of everything will work.
Tortoise TTS vs. Newer Cloning Models
| Tortoise TTS | XTTS-v2 | F5-TTS | |
|---|---|---|---|
| Architecture | Autoregressive transformer + diffusion | Autoregressive transformer (GPT-2-style) + VQ-VAE | Diffusion Transformer + flow matching |
| Generation speed | Slow, quality-first | Moderate | Fast (~0.15 RTF) |
| Voice cloning | Multi-clip reference | Single ~6-second clip | Single ~5–15 second clip |
| Multilingual support | Primarily English | 17 languages | Primarily English-focused |
| License | Apache 2.0 (fully permissive) | Coqui Public Model License (non-commercial) | CC-BY-NC (non-commercial) |
| Best fit | Audiobooks, premium narration, offline quality-first work | General multilingual cloning | Real-time or high-throughput generation |
Tortoise's specific niche today is projects where generation time genuinely doesn't matter and audio quality does — and it remains one of the few well-known cloning models with a license permissive enough to use commercially without a separate agreement.
Frequently Asked Questions
Why is it called "Tortoise TTS"?
The name is a deliberate, tongue-in-cheek reference to how slow the model is — its creator has described it as "insanely slow" by design, prioritizing output quality over generation speed.
Who created Tortoise TTS?
James Betker created Tortoise TTS independently, using his own hardware, with no involvement from any employer — the model and its code and weights are fully open-sourced.
Is Tortoise TTS free for commercial use?
Yes. It's released under the Apache 2.0 license, one of the more permissive options among open-source TTS models, with no restrictions requiring a separate commercial agreement.
How does Tortoise TTS's architecture differ from newer diffusion-based models?
Tortoise pairs an autoregressive transformer (generating discrete speech tokens, similar to how GPT generates text tokens) with a diffusion model that refines those tokens into the final audio — a two-stage design inspired by techniques from image generation, rather than the single flow-matching or diffusion-transformer pipelines used in some newer models.
Did Tortoise TTS influence other TTS models?
Yes. Its combination of an autoregressive language model with a separate refinement stage for audio quality has been cited as an architectural influence on later models, including XTTS and IndexTTS.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.