EchoraEchora
Back to Models

VibeVoice: Built for a Whole Podcast, Not Just a Sentence

September 7, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Almost every TTS model is scoped around a single utterance — generate this sentence, in this voice, and stop. VibeVoice, Microsoft's open-source speech framework, was scoped around a different unit of content entirely: a full episode. It generates up to 90 minutes of continuous, multi-speaker conversational audio in a single pass, with as many as four distinct speakers maintaining consistent identity and natural turn-taking the whole way through — the difference between a model built for a sentence and one built for a podcast.

What Is VibeVoice?

VibeVoice is an open-source, MIT-licensed framework from Microsoft for generating expressive, long-form, multi-speaker conversational audio — podcasts, audio dramas, and interviews are the scenarios it's explicitly built around. It's released in two sizes: VibeVoice-1.5B, built on a Qwen2.5-1.5B language model backbone, and VibeVoice-Large, built on the larger Qwen2.5-7B backbone for higher fidelity at a correspondingly heavier compute cost. Both are hosted openly (microsoft/VibeVoice-1.5B and the Large variant) alongside the project's code on GitHub.

Functional Advantages of VibeVoice

  • 90-minute continuous generation. A single pass can produce up to 90 minutes of coherent audio, maintaining speaker consistency and narrative coherence throughout — far beyond the short-clip scope most TTS models are built for.
  • Up to four distinct speakers, naturally. Multi-speaker dialogue is a first-class capability, not a workaround stitching together separate single-voice generations — turn-taking and speaker identity stay consistent across an entire long-form session.
  • Extremely efficient audio compression. A 7.5Hz tokenization rate achieves roughly 3,200x downsampling from raw audio while preserving fidelity, which is what makes generating 90 minutes of audio computationally tractable in the first place.
  • Emergent cross-lingual transfer and singing. Neither capability was specifically trained for, but both show up as emergent behavior — the model can carry a voice's accent across languages and even attempt singing, though quality on both varies.
  • Responsible-AI safeguards built in by default. Every output includes both an audible disclaimer and an imperceptible watermark, applied automatically rather than as an opt-in setting.
  • Available at multiple scales, including quantized variants. Beyond the base 1.5B and Large checkpoints, community-produced 4-bit and 8-bit quantized versions bring VRAM requirements down substantially for constrained hardware.
  • A companion model for the reverse direction. VibeVoice-ASR extends the same family into speech-to-text, so the project isn't limited to generation alone.

Under the Hood

VibeVoice's pipeline splits into three cooperating pieces. Continuous acoustic and semantic tokenizers, operating at that ultra-low 7.5Hz frame rate, compress raw audio into a representation efficient enough to process genuinely long sequences without the computational cost exploding. A pretrained Qwen2.5 language model (1.5B or 7B, depending on which checkpoint you use) then models the dialogue itself — semantics, flow, and speaker turn structure — conditioning what should be said and by whom. A lightweight four-layer diffusion head paired with a VAE decoder fills in the high-fidelity acoustic detail on top of what the LLM has structured, using classifier-free guidance at inference time to control how closely output adheres to a reference voice's timbre (a cfg_scale of around 1.3 is the commonly recommended setting).

Training is deliberately staged: the tokenizers themselves are frozen, with only the LLM and diffusion head actually trained, using curriculum learning that starts on shorter sequences and progressively extends up to 65,000 tokens — which is specifically what enables the model to hold together coherently over a genuinely long generation rather than drifting or losing speaker consistency partway through.

VibeVoice-ASR: The Listening Half of the Family

Released in January 2026, VibeVoice-ASR extends the same family in the opposite direction — a unified speech-to-text model rather than a text-to-speech one. It processes up to 60 minutes of long-form audio in a single pass, producing structured transcriptions that capture who spoke (through speaker diarization), when (via precise timestamps), and what was said, with support for user-customized context and domain-specific hotwords. It's since been integrated directly into the Hugging Face Transformers library and Microsoft's Azure AI Foundry Labs, making it straightforward to adopt without a bespoke inference setup.

For edge deployment specifically, VibeVoice-ASR-BitNet compresses the model from 4.62GB down to 1.58GB through heterogeneous quantization (I8_S for the audio encoder, I2_S for the language model), swapping the Qwen2.5-7B backbone for a lighter Qwen2.5-1.5B version at only a 1–4% absolute increase in word error rate. Paired with the dedicated VibeASR.cpp runtime — built on custom SIMD kernels and operator fusion within the ggml framework — it delivers real-time transcription (RTF below 1) on as few as 3 CPU threads, with no GPU required, and runs 1.6 to 2.3 times faster than Whisper.cpp at a comparable model size.

Tips for Better Results

  • Script for the format the model actually expects. Structure your input as labeled turns (Speaker 1: ... Speaker 2: ...) in a plain transcript rather than an unstructured paragraph — this is what lets VibeVoice correctly assign turn-taking and maintain per-speaker consistency.
  • Start with the 1.5B model before assuming you need Large. Given the meaningful jump in compute and VRAM requirements, confirm the smaller model isn't already sufficient for your content before committing to the 7B checkpoint.
  • Use a pre-quantized checkpoint for constrained hardware rather than quantizing on the fly. Pre-quantized 4-bit or 8-bit models load faster and avoid the overhead of on-the-fly quantization during every run.
  • Don't rely on emergent singing or cross-lingual output for production content. Since neither was specifically trained or optimized, treat successful results as a bonus rather than a dependable feature, and budget for repeated sampling if you're specifically chasing one of these effects.
  • Keep the audible disclaimer and watermark in mind for your use case. Since both are embedded automatically in every output, factor that into any product documentation around AI-generated audio disclosure.

VibeVoice vs. Other Multi-Speaker and Long-Form Models

VibeVoiceCosyVoice3Typical single-speaker TTS
Max speakers per session4Zero-shot per-clip, not session-based1
Max continuous duration~90 minutesNot a dedicated focusShort clips
Core architectureQwen2.5 LLM + diffusion head + 7.5Hz tokenizersSupervised semantic tokens + flow matchingVaries
Companion ASR modelYes (VibeVoice-ASR)NoNo
Quantized variants availableYes (4-bit/8-bit community + official BitNet for ASR)NoVaries
LicenseMITOpen, hosted API availableVaries

VibeVoice's specific niche is scale of output, not per-sentence expressiveness — if your project is generating a single line of dialogue, most other models on this site will serve you fine; if you're generating an entire podcast episode with multiple hosts, VibeVoice is built specifically for that unit of content.

Frequently Asked Questions

What's the difference between VibeVoice-1.5B and VibeVoice-Large?

1.5B is built on a Qwen2.5-1.5B backbone and is the lighter, faster option; Large uses a Qwen2.5-7B backbone for higher fidelity at a heavier compute and VRAM cost.

How long of an audio clip can VibeVoice actually generate?

Up to 90 minutes of continuous, coherent audio in a single pass, while maintaining consistent speaker identity throughout.

What is VibeVoice-ASR, and is it the same as VibeVoice TTS?

No — VibeVoice-ASR is a separate, companion model in the same family that handles speech-to-text instead of text-to-speech, processing up to 60 minutes of audio with speaker diarization and timestamps.

Can I run VibeVoice-Large on a consumer GPU?

With quantization, yes — pre-quantized 4-bit checkpoints bring VRAM needs down to roughly 10GB, though at some cost to generation speed compared to full precision.

Is VibeVoice free for commercial use?

Yes. It's released under the MIT license, one of the more permissive options among open-source TTS models.

Turn a Multi-Speaker Script into Audio

If you already have a finished multi-speaker script, use Echora’s Text to Dialogue. Add up to 12 dialogue blocks with 5,000 characters in total, assign a voice to each block, add supported delivery tags, and adjust stability or speaker boost. You can then preview and download the complete conversation.

Create Dialogue Audio →