WhisperSpeech: Running a Transcription Model Backwards
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
OpenAI's Whisper is one of the most trusted speech recognition models in existence — trained on an enormous amount of multilingual audio to turn speech into accurate text. Collabora, the team behind WhisperSpeech, looked at that model and asked a genuinely clever question: if Whisper is this good at understanding what speech means, what happens if you run it in reverse? WhisperSpeech is the answer — a text-to-speech system built by inverting Whisper's own architecture rather than designing a TTS pipeline from scratch.
What Is WhisperSpeech?
WhisperSpeech is an open-source TTS system from Collabora, previously developed under the name spear-tts-pytorch before adopting its current name. The project's stated ambition is direct: to be for speech what Stable Diffusion became for images — powerful, hackable, and safe to build commercial products on. That last point is deliberate and specific: all of WhisperSpeech's code is released under Apache-2.0/MIT, and its models are trained exclusively on properly licensed speech data (the English LibreLight dataset for the current release), rather than the murkier, scraped-from-anywhere data provenance that clouds commercial usability for a lot of open TTS research.
How Inverting Whisper Actually Works
Whisper's own architecture splits into two parts: an encoder that turns audio into a continuous stream of embeddings — capturing contextual meaning and prosody along the way — and a decoder that uses cross-attention to convert those embeddings into actual words, i.e., a transcript. WhisperSpeech runs that same structure in the opposite direction: starting from text, it generates the equivalent embeddings, then transforms those embeddings into sound instead of extracting them from sound.
The genuinely clever part is what this reuses. Whisper's encoder, trained on a huge volume of multilingual audio, produces an unusually strong semantic representation of speech content. Architectures like Google's SPEAR-TTS and VALL-E rely on a similar kind of semantic token extraction internally, but their specific semantic encoders were never released publicly. By repurposing Whisper's own encoder for this role, WhisperSpeech provides an open-source, drop-in equivalent to a component that comparable systems had kept closed — which is part of why the project describes itself as filling a real gap rather than just reproducing existing work with new branding.
The Two-Stage Pipeline
WhisperSpeech follows the same broad two-stage, token-based pattern popularized by AudioLM, SPEAR-TTS, and Meta's MusicGen — splitting speech synthesis into a "reading" phase and a "speaking" phase rather than one monolithic step. The text-to-semantic (T2S) stage uses the Whisper-derived approach to convert input text into semantic tokens. The semantic-to-acoustic (S2A) stage then converts those tokens into actual audio, using Meta's EnCodec for acoustic modeling and the Vocos vocoder (from Charactr) to render the final high-quality waveform. Rather than building each of these components from scratch, the project deliberately built on established open-source pieces for every stage — Whisper for semantic understanding, EnCodec for acoustic representation, Vocos for the vocoder — assembling something new from parts that were each already proven independently.
Speed and Emerging Multilingual Capability
After adding torch.compile support, KV-caching, and targeted layer tweaks, WhisperSpeech reached roughly 12 times faster than real-time generation on a consumer RTX 4090 — a meaningful optimization pass that moved it from research-speed to genuinely practical inference speed on accessible hardware. On the multilingual front, the team ran a proof-of-concept: training a small S2A model on a mixed English, Polish, and French dataset, and getting it to successfully clone French voices using semantic tokens that had only ever been trained on English and Polish. That result is a meaningful signal — it suggests a single shared tokenizer could plausibly cover many languages rather than needing a separate one trained per language, though this remains early, exploratory work rather than a finished multilingual release.
Getting Started with WhisperSpeech
The fastest way to hear WhisperSpeech is through the project's provided Google Colab notebook, which runs inference without any local setup. If you'd rather run it yourself, pretrained models and their corresponding converted training datasets are both available on Hugging Face, letting you either run inference locally or use them as a starting point for training your own variant. The project's Discord community, hosted in the LAION server's audio-generation channel, is an active place to ask setup questions or follow ongoing multilingual development.
Tips for Better Results
- Check current language coverage before assuming multilingual support is production-ready. The stable release remains English-focused (via LibreLight); the promising French-cloning result came from a small proof-of-concept model, not the main release — verify current project status before planning around broader language coverage.
- Use the Colab notebook first to validate quality on your content. It's a lower-friction way to hear actual output on your specific use case before committing to local installation and setup.
- Take advantage of the inference speed optimizations if you're deploying at scale. The 12x real-time figure on an RTX 4090 reflects specific optimizations (torch.compile, KV-caching) — make sure your own deployment applies these rather than running an unoptimized baseline configuration.
- Lean on the licensing story if commercial deployment is a concern. Since the project was specifically built around properly licensed training data and permissive code licensing, it's a genuinely lower-risk starting point than TTS projects with less transparent data provenance — worth factoring into any vendor or model selection process where legal review matters.
WhisperSpeech vs. Other Fully Open Cloning Models
| WhisperSpeech | VoiceCraft | F5-TTS | |
|---|---|---|---|
| Core architectural idea | Whisper's ASR pipeline run in reverse | Token infilling for editing and cloning | Diffusion Transformer + flow matching |
| Training data provenance | Exclusively properly licensed (LibreLight) | Open, in-the-wild sources | Emilia dataset (non-commercial license) |
| Code license | Apache-2.0 / MIT | Open | Code open, weights CC-BY-NC |
| Commercial use | Explicitly designed to be safe for this | Open, check ethical use terms | Requires separate clearance |
| Inference speed | ~12x real-time (RTX 4090, optimized) | Not the primary focus | ~0.15 real-time factor |
| Current language coverage | English (multilingual in early testing) | English-focused | Primarily English-focused |
WhisperSpeech's specific niche is the combination of an elegant, resource-efficient architectural idea (reusing Whisper rather than training a new semantic encoder from scratch) with a genuinely deliberate focus on clean data licensing — a combination that matters most for teams where "can we legally ship this" is as important a question as "does it sound good."
Frequently Asked Questions
How does "inverting Whisper" actually produce speech?
Whisper normally converts audio into embeddings and then into text via its encoder-decoder structure. WhisperSpeech reverses that flow — starting from text, generating the equivalent embeddings, and converting those into audio instead of extracting them from it.
What was WhisperSpeech previously called?
It was originally developed under the name spear-tts-pytorch, reflecting its architectural inspiration from Google's SPEAR-TTS research, before being renamed to WhisperSpeech.
Is WhisperSpeech safe to use in a commercial product?
That's an explicit design goal — the code is released under Apache-2.0/MIT, and the models are trained only on properly licensed speech data specifically to avoid the data-provenance risk that affects many open TTS projects trained on scraped or unclear-license audio.
Does WhisperSpeech support languages other than English?
The current stable release is trained on the English LibreLight dataset. A proof-of-concept multilingual model (English, Polish, French) has demonstrated promising cross-lingual voice cloning results, but broader multilingual support was still in progress as of the project's most recent public updates.
How fast is WhisperSpeech at generating audio?
After specific optimizations (torch.compile, KV-caching, and layer adjustments), the project reported roughly 12 times faster than real-time generation on a consumer RTX 4090 GPU.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.