EchoraEchora
Back to Models

Llasa: Speech Synthesis as Just Another Thing LLaMA Learned to Do

September 6, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Llasa, developed by the Audio Lab at the Hong Kong University of Science and Technology, takes a specific bet seriously: if you convert speech into tokens the same way you tokenize text, an existing large language model doesn't need a specialized architecture to generate it — it just needs to learn a new vocabulary. Llasa extends LLaMA models directly, at three different scales, treating speech as what its own research describes as "a special language" rather than a separate modality requiring its own bespoke system.

What Is Llasa?

Llasa is a text-to-speech system that extends the text-based LLaMA language model — available at 1B, 3B, and 8B parameter sizes — by incorporating speech tokens from the XCodec2 codebook, which contains 65,536 distinct tokens. It was trained on 250,000 hours of Chinese-English speech data, combining roughly 160,000 hours from open-source corpora (LibriHeavy, Emilia, and WenetSpeech4TTS) with an additional 90,000 hours of internal data. The method is built to be seamlessly compatible with the standard LLaMA framework — audio is converted into single-codebook tokens and treated as just another sequence a language model can learn to predict, which means existing techniques for LLM compression, acceleration, and fine-tuning can be applied to Llasa directly, rather than needing TTS-specific tooling built from scratch.

Choosing Between 1B, 3B, and 8B

Having three sizes of the same underlying approach is a genuine practical advantage, letting you match model scale to your actual hardware and quality needs rather than being locked into a single fixed footprint:

  • Llasa-1B is the lightest option, the most practical choice when resource constraints matter more than squeezing out maximum quality, and the version most commonly used as a base for further fine-tuning on new languages or dialects.
  • Llasa-3B sits in the middle, and is the version most frequently referenced in community demos and comparisons — a reasonable default if you're unsure which size fits your use case.
  • Llasa-8B is the largest and most capable version, extending further into complex text comprehension and nuanced speech generation, at a correspondingly larger compute cost.

Scaling Test-Time Compute for Speech

This is Llasa's most distinctive research contribution, and it borrows directly from a trend that's driven progress in text-based LLMs recently: rather than only improving quality by training a bigger model or on more data, Llasa's research explores scaling compute at inference time as well — the same broad idea behind reasoning models that "think longer" to produce better answers, applied here to speech generation instead of text reasoning. This positions Llasa's contribution as much about how to get more out of a fixed model through inference-time strategies as it is about the model architecture itself.

Strong Text Comprehension

Because Llasa is built directly on a genuine language model foundation rather than a text-to-phoneme pipeline bolted onto an acoustic model, it handles structurally complex text with a level of comprehension that trips up many dedicated TTS systems — numbered lists, embedded punctuation, and multi-clause sentences with quoted dialogue are specifically highlighted in the project's own examples as text Llasa-8B navigates correctly rather than stumbling over.

Voice Cloning via Speech Prompts

Llasa generates speech either purely from input text, using its own general voice characteristics, or by conditioning on a speech prompt for voice cloning — reported to require as little as 15 seconds of reference audio to capture a target speaker's timbre and emotional character. This dual-mode design means you're not forced to clone a specific voice if you don't need to; plain text-to-speech is available as a first-class mode rather than an afterthought.

A Practical Quirk Worth Knowing

If you're working with Llasa-3B specifically, note that the model expects audio input at 16kHz — feeding in audio at a different sample rate is a documented source of degraded results, so resampling your reference clips to match before running inference is worth doing as a standard preprocessing step rather than discovering the mismatch after a disappointing generation.

Getting Started with Llasa

  1. Install XCodec2 first. Since Llasa's speech tokenization depends on it directly, get this dependency working before attempting any inference or training.
  2. Choose your model size from Hugging Face. HKUSTAudio/Llasa-1B, HKUSTAudio/Llasa-3B, and HKUSTAudio/Llasa-8B are all available as separate downloads — pick based on your hardware budget and quality needs rather than defaulting to the largest available.
  3. Use the dedicated inference repository for testing. The zhenye234/LLaSA_inference GitHub project is built specifically for running and experimenting with test-time compute scaling, separate from the training codebase.
  4. Use the training repository if you're fine-tuning or training from scratch. zhenye234/LLaSA_training provides the scripts and documented open-source tokenized datasets (roughly 160,000 hours across LibriHeavy, Emilia, and WenetSpeech4TTS) needed to reproduce or extend the base training process.
  5. Match your reference audio sample rate to the model's expectations. Particularly for Llasa-3B's documented 16kHz requirement, resample your speech prompt audio before generation rather than assuming any sample rate will work interchangeably.
  6. Consider fine-tuning Llasa-1B for languages or dialects beyond its base training. Since the smaller model is the more common fine-tuning starting point, and given documented community work extending it to additional languages like Cantonese, this is a realistic path if your target language isn't well covered by the base Chinese-English training.

Tips for Better Results

  • Start with Llasa-3B if you're unsure which size to test first. It's the most commonly referenced version in community usage and demos, making it easier to find comparison points and troubleshooting help than with the less-discussed 1B or 8B checkpoints.
  • Feed it genuinely complex text to see its real strength. Since strong text comprehension is a specifically highlighted capability, testing with simple, clean sentences alone won't show you where Llasa actually differentiates itself from more basic TTS pipelines.
  • Use a 15-second-or-longer clean reference clip for cloning. This is the practical minimum commonly cited for capturing a target voice's timbre and emotional character reliably.
  • Don't expect strong out-of-the-box performance on languages outside its Chinese-English training. Independent evaluation on Cantonese, for instance, showed the base Llasa-1B model performing meaningfully worse than dedicated multilingual models until it was specifically fine-tuned on Cantonese data — plan for a fine-tuning step if your target language isn't Chinese or English.
  • Leverage LLM-style techniques given the shared LLaMA architecture. Because Llasa is genuinely built on the LLaMA framework, established LLM optimization approaches — quantization, LoRA fine-tuning, and similar techniques — are more likely to transfer cleanly than they would to a fundamentally different TTS architecture.

Llasa vs. Other LLM-Based TTS Approaches

LlasaOuteTTSCosyVoice3
Base architectureLLaMA (1B/3B/8B) + XCodec2 tokensLLaMA/Qwen + DAC tokensSupervised semantic tokens + flow matching
Model size options1B, 3B, 8B350M–1B0.5B
LanguagesChinese, English23 (v1.0)9 + 18 Chinese dialects
Voice cloningYes, ~15 second speech promptYes, ~10 second referenceYes, ~3–10 second reference
Runs via llama.cpp / GGUFNot the primary deployment pathYes, nativelyNo
Out-of-the-box dialect performanceRequires fine-tuning beyond Chinese/EnglishEnglish-focusedStrong native Chinese dialect support

Llasa's specific niche is offering genuine model-size flexibility (1B through 8B) within the "TTS as pure language modeling" approach, paired with a research focus on scaling inference-time compute — a different emphasis than OuteTTS's llama.cpp-first deployment story or CosyVoice3's dialect-breadth focus, even though all three sit in a similar architectural family.

Frequently Asked Questions

What's the difference between Llasa 1B, 3B, and 8B?

They're the same underlying approach at increasing scale — 1B is the lightest and most common fine-tuning starting point, 3B is the most frequently used version in community demos, and 8B is the largest, offering the strongest text comprehension and generation quality at a higher compute cost.

How much reference audio does Llasa need for voice cloning?

Around 15 seconds of clean reference audio is commonly cited as enough to capture a target speaker's timbre and emotional character.

Does Llasa work well in languages other than Chinese and English?

Not out of the box. Independent testing on Cantonese, for example, showed the base model underperforming dedicated Cantonese or multilingual systems until it was specifically fine-tuned on that language's data.

What does "scaling test-time compute" mean for Llasa?

It refers to improving output quality by using additional compute at inference time, rather than only through training-time scale — a similar concept to how reasoning-focused language models use extra inference compute to produce better answers, applied here to speech generation.

Why does Llasa-3B specifically need 16kHz audio input?

This is a documented requirement of that particular checkpoint — feeding in audio at other sample rates is a known source of degraded results, so resampling reference clips to 16kHz beforehand is worth doing as standard practice.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →