EchoraEchora
Back to Models

Qwen3-TTS: A Voice Model That Talks Back in Under a Tenth of a Second

August 29, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Most voice-cloning demos make you wait a beat before the audio starts — long enough that it never quite feels like a conversation. Qwen3-TTS, the open-source speech model from Alibaba's Qwen team, was built to close that gap: a dual-track architecture that generates end-to-end speech with a first-audio-packet latency as low as 97 milliseconds, fast enough to sit behind a live voice agent without adding a noticeable delay of its own.

What Is Qwen3-TTS?

Qwen3-TTS is an open-source text-to-speech model trained on more than 5 million hours of speech data, released under the Apache 2.0 license with public weights on Hugging Face and GitHub. At its core is a custom speech tokenizer, Qwen3-TTS-Tokenizer-12Hz, paired with a hybrid streaming architecture that generates audio directly rather than routing through a separate diffusion or vocoder stage — the design choice behind both its low latency and its stability on longer or noisier input text.

The model covers 10 major languages — Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian — along with several Chinese dialectal voice profiles, and ships in two open-weight sizes: a 1.7B-parameter version tuned for maximum quality, and a lighter 0.6B version built for constrained hardware without giving up much in return.

Three Ways to Get a Voice Out of It

Qwen3-TTS doesn't lock you into one workflow — it's built around three distinct ways to generate speech, depending on what you actually have to work with.

Clone a voice from 3 seconds of audio. Provide a short reference clip along with its matching transcript, and the model reproduces that speaker's voice on new text — one of the shortest reference-audio requirements among current open-source cloning models.

Design a voice from a text description. No recording on hand? Describe the voice in plain language — age, gender, tone, personality — and Qwen3-TTS generates a new voice to match, without needing any audio reference at all. It's a genuinely useful option when you need a distinct character voice and don't want to source, license, or record one first.

Pick a ready-made preset. For projects that just need a solid voice without any setup, 49 preset multi-character voices are available through the hosted API, built around named personas covering a range of ages, genders, and personalities, swappable instantly with no additional configuration.

Across all three modes, the model reads context well — it adjusts tone, pacing, and emotional delivery based on the semantics of the text itself, and holds up better against messy or imperfectly formatted input than earlier speech models in the same family.

Getting Started with Qwen3-TTS

  1. Decide between self-hosting and the API. Self-hosting the open weights gives you full control and no per-character cost; calling the hosted API (model names like qwen3-tts-flash for fast, low-latency generation, or qwen3-tts-instruct-flash when you need natural-language instruction control) is the quicker path if you'd rather not manage GPU infrastructure.
  2. Install the package. pip install qwen-tts gives you the Qwen3TTSModel class needed to load and run inference locally.
  3. Pick a checkpoint sized to your hardware. Qwen/Qwen3-TTS-12Hz-1.7B-Base is the full-quality model, needing roughly 6–8GB of VRAM; the 0.6B variant runs comfortably on 4–6GB with a smaller trade-off in quality than the size difference suggests.
  4. Generate a cloned voice. Call generate_voice_clone() with your target text, a language code, a reference audio file, and its matching transcript.
  5. Cache reference prompts for repeated generations. If you're producing multiple clips from the same cloned voice, reuse the computed prompt features instead of recalculating them on every call — it noticeably speeds up batch generation.
  6. Use vLLM for production throughput. The Qwen team supports vLLM for Qwen3-TTS from day one, which is worth using over a plain PyTorch inference loop for anything beyond small-scale testing.

Trying It Without Setting Anything Up

The official Hugging Face Space and ModelScope demo both let you test voice cloning and voice design directly in the browser, no installation required. Alibaba Cloud has also offered free character quotas for developers trying the model through its Model Studio API, and third-party platforms like Replicate and DeepInfra host the open-weight model as well, which is a low-friction way to evaluate quality before committing to either local infrastructure or a paid API tier.

Tips for Better Results

  • Get the reference transcript exactly right for cloning. Cloning quality depends heavily on the reference text matching what's actually said in the audio — a mismatched transcript hurts results more than a slightly lower-quality recording would.
  • Reach for voice design when a real reference clip isn't available. Describing the voice you want often gets you closer to the result you need than searching for a stock clip that's "close enough," and sidesteps any licensing questions around a real recorded voice.
  • Start with the 0.6B model if you're unsure about hardware. At about half the footprint of the 1.7B version, it holds up well for most use cases — move up only if you specifically need the extra fidelity.
  • Use the hosted, low-latency model variants for live applications. If you're building something closer to a voice agent than batch content generation, the API's fast, low-latency options are built specifically for that use case, so you don't have to optimize self-hosted inference for it yourself.

Qwen3-TTS vs. Other Open-Source Cloning Models

Qwen3-TTSCosyVoice3Chatterbox
Backing organizationAlibaba (Qwen team)Alibaba (FunAudioLLM)Resemble AI
Parameters0.6B / 1.7B0.5B~0.5B
Reference clip for cloning~3 seconds~3–10 seconds5–10 seconds
First-packet latency~97ms~150ms (hosted)Not a primary focus
Voice design from textYesNoNo
Languages109 + 18 Chinese dialectsEnglish (Multilingual variant: 23)
LicenseApache 2.0Open, hosted API availableMIT

Qwen3-TTS's specific edge is combining genuinely low latency with three distinct generation paths — cloning, text-based voice design, and ready-made presets — in a single open-weight release, a broader feature set than most single-purpose cloning models offer at a comparable size.

Frequently Asked Questions

Is Qwen-TTS the same model as Qwen3-TTS?

Not quite. Qwen-TTS was Alibaba's earlier speech model, available only through the hosted Qwen API with no public weights. Qwen3-TTS is the newer generation, open-sourced with downloadable weights and a broader feature set — if you're looking to self-host or fine-tune, Qwen3-TTS is the one to use.

What does qwen3-tts-flash mean?

It's the name of a hosted, managed version of Qwen3-TTS available through Alibaba Cloud's API, tuned for fast, low-latency generation without needing to self-host. It runs the same underlying model family as the open-weight release, just as a managed service.

Is Qwen3-TTS free to use?

The open weights are free to self-host under the Apache 2.0 license. The hosted API has offered free character quotas for new developers, with paid usage beyond that — check current Alibaba Cloud pricing for exact limits.

What's the practical difference between the 0.6B and 1.7B models?

The 1.7B model gives you higher output quality and needs more VRAM (roughly 6–8GB); the 0.6B model is lighter (4–6GB) with a quality trade-off that's smaller than the parameter gap would suggest.

Can I test Qwen3-TTS before installing anything?

Yes — the official Hugging Face Space and ModelScope demo both run in-browser, and hosting platforms like Replicate and DeepInfra offer API-based access without local setup.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.

Create a Voice Clone →