EchoraEchora
Back to Models

Spark-TTS: One Codec, No Separate Acoustic Model Required

September 12, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Most LLM-based TTS systems need a second model after the language model does its job — a flow-matching network or diffusion decoder to turn predicted tokens into actual acoustic detail. Spark-TTS, developed collaboratively by researchers at HKUST, Mobvoi, Shanghai Jiao Tong University, and several other institutions, removes that second stage entirely. Built on Qwen2.5, it reconstructs audio directly from the codes the language model predicts, using a codec specifically designed to make that direct reconstruction possible without sacrificing quality or control.

What Is Spark-TTS?

Spark-TTS is an open-source, LLM-based text-to-speech system introduced in a March 2025 paper, built entirely on the Qwen2.5 language model as its backbone. Its central contribution is BiCodec, a single-stream speech codec that eliminates the need for the separate acoustic-generation models — like flow matching — that most comparable LLM-based TTS architectures rely on. Rather than predicting multiple stacked codebooks that then require a dedicated decoder to interpret, Spark-TTS's LLM predicts codes that map directly to audio, streamlining the entire generation pipeline into something simpler and more computationally efficient.

Functional Advantages of Spark-TTS

  • A genuinely simplified architecture. No separate flow-matching or diffusion model is needed to generate acoustic features — audio is reconstructed directly from LLM-predicted codes, reducing both complexity and inference cost.
  • Zero-shot voice cloning, replicating a target speaker's voice without any dedicated training data for that specific voice.
  • A separate Voice Creation mode, generating an entirely synthetic voice from control parameters alone, with no reference audio required at all — a genuinely distinct capability from cloning.
  • Fine-grained control over delivery, including precise pitch values and speaking rate, alongside coarser controls like gender and general speaking style.
  • Bilingual Chinese-English support, including cross-lingual code-switching scenarios within the same generation.
  • A fully open research pipeline, with the model, training code, and the purpose-built VoxBox dataset (100,000 hours of speech with detailed attribute annotations) all released publicly.

How BiCodec's Disentangled Tokens Actually Work

This is the specific design choice that makes Spark-TTS's simplicity possible without giving up control, and it's worth understanding directly. BiCodec decomposes speech into two complementary token types: low-bitrate semantic tokens that capture linguistic content (what's being said), and fixed-length global tokens that capture speaker-specific attributes (who's saying it, and in what style). Because these two token types are cleanly separated rather than entangled together, the Qwen2.5 backbone — paired with a chain-of-thought generation approach — can manipulate them somewhat independently: adjusting the global tokens shifts voice characteristics without disrupting linguistic content, and vice versa. This disentanglement is specifically what enables both coarse-grained control (picking a gender or general speaking style) and fine-grained control (dialing in an exact pitch value or speaking rate) from the same underlying representation, rather than needing separate mechanisms for each level of control.

Two Modes: Voice Cloning and Voice Creation

Spark-TTS is built around two genuinely distinct workflows rather than one. Voice Cloning takes a reference audio clip and its transcript, and generates new speech in that speaker's voice — the standard zero-shot cloning use case. Voice Creation works the other way around: instead of a reference recording, you specify control parameters (gender, pitch, speaking rate, style) and the model generates an entirely synthetic voice matching those specifications, with no real person's voice involved at all. This second mode is genuinely useful for projects that need a consistent virtual speaker identity — a branded assistant voice, a game character — without needing to source, license, or record a real voice actor as the starting reference.

Getting Started with Spark-TTS

  1. Clone the repository and set up a Python 3.12 environment. git clone https://github.com/SparkAudio/Spark-TTS.git, then create and activate a dedicated conda environment (conda create -n sparktts -y python=3.12 && conda activate sparktts) before installing dependencies with pip install -r requirements.txt.
  2. Download the model weights. Use snapshot_download("SparkAudio/Spark-TTS-0.5B", local_dir="pretrained_models/Spark-TTS-0.5B") from the huggingface_hub library, or clone directly via git lfs install && git clone https://huggingface.co/SparkAudio/Spark-TTS-0.5B pretrained_models/Spark-TTS-0.5B.
  3. Run cloning inference via the CLI. python -m cli.inference --text "text to synthesize." --device 0 --save_dir "path/to/save/audio" --model_dir pretrained_models/Spark-TTS-0.5B --prompt_text "transcript of the prompt audio" --prompt_speech_path "path/to/prompt_audio" covers the basic voice-cloning workflow directly from the command line.
  4. Launch the WebUI for both Voice Cloning and Voice Creation. python webui.py --device 0 opens a browser interface supporting either workflow — Voice Cloning accepts an uploaded reference clip or a live recording, while Voice Creation exposes the control parameters directly.
  5. Look at the Triton/TensorRT-LLM deployment reference for production serving. The project documents an Nvidia Triton and TensorRT-LLM deployment path specifically for teams moving beyond local experimentation into production-scale serving.

Tips for Better Results

  • Use Voice Creation instead of cloning when you don't need a specific real person's voice. If your project just needs a consistent, distinct virtual speaker, generating one through control parameters avoids the licensing and consent questions that come with cloning a real reference voice.
  • Provide an accurate prompt transcript for cloning. Since the CLI requires both the reference audio and its matching transcript text, a precise transcript improves how reliably the model separates linguistic content from speaker identity during generation.
  • Experiment with fine-grained pitch and rate controls rather than relying on coarse settings alone. Given that BiCodec's disentangled tokens specifically enable precise adjustment, using exact values rather than only broad style categories is where the model's control genuinely differentiates itself from reference-only cloning systems.
  • Test cross-lingual code-switching directly if your content needs it. Since this is an explicitly supported scenario rather than an edge case, validating your specific Chinese-English mixed content directly is worthwhile before assuming uniform quality.

Spark-TTS vs. Other LLM-Based Cloning Models

Spark-TTSCosyVoice3F5-TTS
Core mechanismSingle-stream BiCodec, direct LLM-to-audioSupervised semantic tokens + flow matchingDiffusion Transformer + flow matching
Separate acoustic model neededNoYes (flow matching stage)Yes (the core generation stage)
Zero-shot cloningYesYesYes
Fully synthetic voice creation (no reference)Yes, dedicated modeNot a dedicated featureNot a dedicated feature
Fine-grained pitch/rate controlYes, via disentangled tokensInstruction-basedNot a dedicated feature
LanguagesChinese, English (code-switching)9 + 18 Chinese dialectsPrimarily English-focused
LicenseCC-BY-NC-SA-4.0 (non-commercial, share-alike)Open, hosted API availableCC-BY-NC (non-commercial)

Spark-TTS's specific niche is architectural simplicity paired with genuine dual-purpose flexibility — the single-stream design cuts out a whole model stage that most comparable systems still need, while its Voice Creation mode covers a use case (a synthetic voice with no real reference at all) that most cloning-focused models don't address directly.

Frequently Asked Questions

What makes Spark-TTS's architecture different from other LLM-based TTS models?

It uses BiCodec, a single-stream codec that lets the Qwen2.5 backbone predict codes that map directly to audio, eliminating the separate flow-matching or diffusion model that most comparable systems need as an additional generation stage.

Can Spark-TTS create a voice without any reference audio?

Yes — that's the specific purpose of its Voice Creation mode, which generates a fully synthetic voice from control parameters like gender, pitch, and speaking rate, with no reference recording required at all.

Does Spark-TTS support languages other than English and Chinese?

Its documented and trained focus is bilingual Chinese-English, including cross-lingual code-switching within the same generation — broader language support isn't part of the core release.

Is Spark-TTS free for commercial use?

Not without separate clearance. It's released under a CC-BY-NC-SA-4.0 license, restricting use to non-commercial purposes with share-alike terms, and its documentation explicitly limits intended use to research, education, and legitimate applications while prohibiting unauthorized cloning or impersonation.

How much reference audio does Spark-TTS need for voice cloning?

The documented workflow uses a short reference clip along with its matching transcript text; exact minimum duration isn't specified as rigidly as in some other models, but a clean, representative sample is what the cloning pipeline is built around.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar voice-cloning goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate speech that aims to resemble the reference voice.

Create a Voice Clone →