OuteTTS: A Voice Model That's Secretly Just an LLM
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most TTS models need their own dedicated inference pipeline — a specialized runtime that knows how to handle vocoders, mel-spectrograms, and audio-specific architecture quirks. OuteTTS, from OuteAI, sidesteps that entirely by treating speech generation as nothing more than language modeling: predict the next token, except the tokens represent audio instead of words. Because that's genuinely all it's doing under the hood, OuteTTS runs directly through llama.cpp — the same lightweight engine that runs any other GGUF language model — with no TTS-specific runtime required.
What Is OuteTTS?
OuteTTS is an experimental text-to-speech project built entirely on standard large language model architecture, using crafted prompts and audio tokens rather than a specialized TTS pipeline with external adapters or vocoder-specific components bolted on. Early versions were built on a custom LLaMA-based foundation (Oute3-350M-DEV); later releases moved to established backbones — Qwen-2.5-0.5B for the 0.2 release, and a Llama architecture for the current flagship, Llama-OuteTTS-1.0-1B. The through-line across every version is the same: if you can run a language model, you can run OuteTTS, because architecturally, that's exactly what it is.
The Version History
- OuteTTS-0.1-350M — the original proof of concept, demonstrating that pure language modeling could produce usable speech synthesis without a bespoke TTS architecture. At 350M parameters, it's compact enough to run almost anywhere, but the tradeoff shows: the model can frequently alter, insert, or omit words, leading to noticeably inconsistent output quality.
- OuteTTS-0.2-500M — built on the Qwen-2.5-0.5B foundation and trained on more than 5 billion audio prompt tokens, adding a DAC (Descript Audio Codec) audio encoder with two codebooks for meaningfully better audio reconstruction quality, alongside improved voice cloning.
- OuteTTS-0.3-500M — an intermediate refinement with prompt-format changes, which is also why compatibility across community tools and llama.cpp integrations sometimes has to be checked version by version.
- Llama-OuteTTS-1.0-1B — the current flagship release, adding automatic word alignment (raw text input, no preprocessing needed) and meaningfully expanded multilingual coverage, while keeping the overall footprint compact at 1 billion parameters.
Why llama.cpp Specifically
This isn't an arbitrary compatibility choice — it's a direct consequence of OuteTTS's pure-language-modeling design, and there's a specific technical reason llama.cpp in particular produces the most reliable results among the backends OuteTTS supports.
OuteTTS 1.0's sampling requires repetition penalty to be applied to a recent 64-token window, rather than across the entire context window — penalizing the full context causes broken or degraded output. llama.cpp handles this windowed penalty correctly by default, which is part of why it's historically delivered the most consistent output quality among supported backends. Other backends, including standard Hugging Face Transformers, didn't originally implement this windowed approach natively; the outetts Python library has since added a patch replicating the same windowed penalty for the Transformers backend specifically to close that quality gap, but llama.cpp's native handling remains the more dependable default.
Beyond that specific sampling detail, running as a standard GGUF model also means OuteTTS inherits everything the broader llama.cpp ecosystem already provides: quantization options for smaller memory footprints, GPU layer offloading via n_gpu_layers, and compatibility with tools already built around GGUF-format language models, rather than needing a TTS-specific deployment story built from scratch.
Voice Cloning and Multilingual Support
Llama-OuteTTS-1.0-1B supports one-shot voice cloning from just 10 seconds of reference audio, and covers 23 languages with the model handling text alignment automatically — including languages without clear word boundaries, where automatic alignment matters more than in space-delimited languages like English. Audio reconstruction runs through the DAC encoder inherited from earlier versions, which is part of why quality held up as multilingual coverage expanded rather than degrading as more languages were added.
Getting Started with OuteTTS
- Install the Python library.
pip install outettsgets you the core interface; if you're specifically using GGUF models through llama.cpp, you'll also need to manually installllama-cpp-pythonfirst, following its own installation instructions for your platform. - Use the automatic configuration helper for the simplest setup.
outetts.ModelConfig.auto_config(model=outetts.Models.VERSION_1_0_SIZE_1B, backend=outetts.Backend.LLAMACPP, quantization=outetts.LlamaCppQuantization.FP16)handles model selection and backend configuration without manual path management. - Configure manually if you need specific file paths. A manual
ModelConfigneeds bothmodel_path(pointing to your downloaded.gguffile) andtokenizer_path(which is always Transformers-based, even when your inference backend is llama.cpp) — mixing these up is a common source of loading errors. - Download GGUF weights from the dedicated repository. GGUF-format checkpoints are hosted separately from the standard safetensors weights — make sure you're pulling from the
-GGUFrepository variant specifically if you're using the llama.cpp backend. - Set
n_gpu_layersif you want GPU acceleration. As with any llama.cpp-based model, offloading layers to GPU is controlled through this standard parameter rather than anything OuteTTS-specific. - Let the
outettslibrary handle sampling configuration automatically. Given how specific and consequential the windowed repetition penalty requirement is, rely on the library's built-in sampler setup rather than reimplementing generation parameters yourself unless you have a clear reason to.
Tips for Better Results
- Default to Llama-OuteTTS-1.0-1B rather than an older version. Given the word-accuracy issues documented in the original 350M release and the prompt-format inconsistencies across 0.2 and 0.3, the current 1.0 release is the more dependable starting point for new projects.
- Trust llama.cpp's default sampling behavior over a custom implementation. Unless you have a specific reason to build your own inference wrapper, the windowed repetition penalty that llama.cpp handles correctly by default is easy to get wrong in a from-scratch implementation.
- Use a clean 10-second reference clip for cloning. Since one-shot cloning is designed around that specific duration, a well-recorded clip at roughly that length is more likely to produce reliable results than either a much shorter or unnecessarily longer sample.
- Feed in raw text for version 1.0 rather than pre-processing it yourself. Since automatic word alignment is a specific feature of the current release, manually pre-segmenting your text works against a capability the model already handles internally.
- Check the license before any commercial use. OuteTTS's GGUF releases are distributed under a CC-BY-NC-SA-4.0 license — non-commercial, share-alike — which is a meaningfully different set of terms than a permissive license like MIT or Apache 2.0. Review this carefully before building a commercial product on it.
OuteTTS vs. Other Lightweight, LLM-Adjacent TTS Models
| OuteTTS | Piper | Kokoro-82M | |
|---|---|---|---|
| Core approach | Pure language modeling on audio tokens | VITS + ONNX + espeak-ng | StyleTTS 2 + ISTFTNet |
| Runs on llama.cpp | Yes, natively (GGUF) | No | No |
| Voice cloning | Yes, one-shot from ~10 seconds | No, fixed voice roster | No, fixed voice roster |
| Languages | 23 | 35+ | 8 |
| Parameters | 350M–1B depending on version | ~15M | 82M |
| License | CC-BY-NC-SA-4.0 (non-commercial) | MIT (original) / GPL-3.0 (current fork) | Apache 2.0 |
OuteTTS's specific niche is architectural simplicity for anyone already working in the LLM tooling ecosystem — if your infrastructure is already built around llama.cpp and GGUF models, adding voice generation doesn't require a separate deployment stack the way most other TTS options would.
Frequently Asked Questions
Why does OuteTTS work with llama.cpp when most TTS models don't?
Because OuteTTS is built as pure language modeling — it predicts audio tokens the same way a standard LLM predicts text tokens, with no specialized TTS architecture layered on top. That means it's compatible with any tool built to run GGUF-format language models, llama.cpp included.
How much reference audio does OuteTTS need for voice cloning?
Around 10 seconds for one-shot cloning with the current Llama-OuteTTS-1.0-1B release.
Why does my custom OuteTTS implementation produce broken audio?
The most common cause is an incorrectly configured repetition penalty. OuteTTS 1.0 requires this penalty to be applied to a recent 64-token window rather than the full context — llama.cpp handles this correctly by default, but a custom implementation needs to replicate that windowing explicitly.
Is OuteTTS free for commercial use?
Not without review. OuteTTS's GGUF releases are distributed under a CC-BY-NC-SA-4.0 license, which restricts commercial use and requires share-alike licensing for derivatives — check current terms before any commercial deployment.
Which OuteTTS version should I use?
Llama-OuteTTS-1.0-1B is the current flagship and the more dependable choice for new projects, offering automatic word alignment and broader multilingual support compared to the earlier 0.1, 0.2, and 0.3 releases.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.