EchoraEchora
Back to Models

WhisperSpeech: Running a Transcription Model Backwards

September 8, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

OpenAI's Whisper is one of the most trusted speech recognition models in existence — trained on an enormous amount of multilingual audio to turn speech into accurate text. Collabora, the team behind WhisperSpeech, looked at that model and asked a genuinely clever question: if Whisper is this good at understanding what speech means, what happens if you run it in reverse? WhisperSpeech is the answer — a text-to-speech system built by inverting Whisper's own architecture rather than designing a TTS pipeline from scratch.

What Is WhisperSpeech?

WhisperSpeech is an open-source TTS system from Collabora, previously developed under the name spear-tts-pytorch before adopting its current name. The project's stated ambition is direct: to be for speech what Stable Diffusion became for images — powerful, hackable, and safe to build commercial products on. That last point is deliberate and specific: all of WhisperSpeech's code is released under Apache-2.0/MIT, and its models are trained exclusively on properly licensed speech data (the English LibreLight dataset for the current release), rather than the murkier, scraped-from-anywhere data provenance that clouds commercial usability for a lot of open TTS research.

How Inverting Whisper Actually Works

Whisper's own architecture splits into two parts: an encoder that turns audio into a continuous stream of embeddings — capturing contextual meaning and prosody along the way — and a decoder that uses cross-attention to convert those embeddings into actual words, i.e., a transcript. WhisperSpeech runs that same structure in the opposite direction: starting from text, it generates the equivalent embeddings, then transforms those embeddings into sound instead of extracting them from sound.

The genuinely clever part is what this reuses. Whisper's encoder, trained on a huge volume of multilingual audio, produces an unusually strong semantic representation of speech content. Architectures like Google's SPEAR-TTS and VALL-E rely on a similar kind of semantic token extraction internally, but their specific semantic encoders were never released publicly. By repurposing Whisper's own encoder for this role, WhisperSpeech provides an open-source, drop-in equivalent to a component that comparable systems had kept closed — which is part of why the project describes itself as filling a real gap rather than just reproducing existing work with new branding.

The Two-Stage Pipeline

WhisperSpeech follows the same broad two-stage, token-based pattern popularized by AudioLM, SPEAR-TTS, and Meta's MusicGen — splitting speech synthesis into a "reading" phase and a "speaking" phase rather than one monolithic step. The text-to-semantic (T2S) stage uses the Whisper-derived approach to convert input text into semantic tokens. The semantic-to-acoustic (S2A) stage then converts those tokens into actual audio, using Meta's EnCodec for acoustic modeling and the Vocos vocoder (from Charactr) to render the final high-quality waveform. Rather than building each of these components from scratch, the project deliberately built on established open-source pieces for every stage — Whisper for semantic understanding, EnCodec for acoustic representation, Vocos for the vocoder — assembling something new from parts that were each already proven independently.

Speed and Emerging Multilingual Capability

After adding torch.compile support, KV-caching, and targeted layer tweaks, WhisperSpeech reached roughly 12 times faster than real-time generation on a consumer RTX 4090 — a meaningful optimization pass that moved it from research-speed to genuinely practical inference speed on accessible hardware. On the multilingual front, the team ran a proof-of-concept: training a small S2A model on a mixed English, Polish, and French dataset, and getting it to successfully clone French voices using semantic tokens that had only ever been trained on English and Polish. That result is a meaningful signal — it suggests a single shared tokenizer could plausibly cover many languages rather than needing a separate one trained per language, though this remains early, exploratory work rather than a finished multilingual release.

Getting Started with WhisperSpeech

The fastest way to hear WhisperSpeech is through the project's provided Google Colab notebook, which runs inference without any local setup. If you'd rather run it yourself, pretrained models and their corresponding converted training datasets are both available on Hugging Face, letting you either run inference locally or use them as a starting point for training your own variant. The project's Discord community, hosted in the LAION server's audio-generation channel, is an active place to ask setup questions or follow ongoing multilingual development.

Tips for Better Results

  • Check current language coverage before assuming multilingual support is production-ready. The stable release remains English-focused (via LibreLight); the promising French-cloning result came from a small proof-of-concept model, not the main release — verify current project status before planning around broader language coverage.
  • Use the Colab notebook first to validate quality on your content. It's a lower-friction way to hear actual output on your specific use case before committing to local installation and setup.
  • Take advantage of the inference speed optimizations if you're deploying at scale. The 12x real-time figure on an RTX 4090 reflects specific optimizations (torch.compile, KV-caching) — make sure your own deployment applies these rather than running an unoptimized baseline configuration.
  • Lean on the licensing story if commercial deployment is a concern. Since the project was specifically built around properly licensed training data and permissive code licensing, it's a genuinely lower-risk starting point than TTS projects with less transparent data provenance — worth factoring into any vendor or model selection process where legal review matters.

WhisperSpeech vs. Other Fully Open Cloning Models

WhisperSpeechVoiceCraftF5-TTS
Core architectural ideaWhisper's ASR pipeline run in reverseToken infilling for editing and cloningDiffusion Transformer + flow matching
Training data provenanceExclusively properly licensed (LibreLight)Open, in-the-wild sourcesEmilia dataset (non-commercial license)
Code licenseApache-2.0 / MITOpenCode open, weights CC-BY-NC
Commercial useExplicitly designed to be safe for thisOpen, check ethical use termsRequires separate clearance
Inference speed~12x real-time (RTX 4090, optimized)Not the primary focus~0.15 real-time factor
Current language coverageEnglish (multilingual in early testing)English-focusedPrimarily English-focused

WhisperSpeech's specific niche is the combination of an elegant, resource-efficient architectural idea (reusing Whisper rather than training a new semantic encoder from scratch) with a genuinely deliberate focus on clean data licensing — a combination that matters most for teams where "can we legally ship this" is as important a question as "does it sound good."

Frequently Asked Questions

How does "inverting Whisper" actually produce speech?

Whisper normally converts audio into embeddings and then into text via its encoder-decoder structure. WhisperSpeech reverses that flow — starting from text, generating the equivalent embeddings, and converting those into audio instead of extracting them from it.

What was WhisperSpeech previously called?

It was originally developed under the name spear-tts-pytorch, reflecting its architectural inspiration from Google's SPEAR-TTS research, before being renamed to WhisperSpeech.

Is WhisperSpeech safe to use in a commercial product?

That's an explicit design goal — the code is released under Apache-2.0/MIT, and the models are trained only on properly licensed speech data specifically to avoid the data-provenance risk that affects many open TTS projects trained on scraped or unclear-license audio.

Does WhisperSpeech support languages other than English?

The current stable release is trained on the English LibreLight dataset. A proof-of-concept multilingual model (English, Polish, French) has demonstrated promising cross-lingual voice cloning results, but broader multilingual support was still in progress as of the project's most recent public updates.

How fast is WhisperSpeech at generating audio?

After specific optimizations (torch.compile, KV-caching, and layer adjustments), the project reported roughly 12 times faster than real-time generation on a consumer RTX 4090 GPU.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →