EchoraEchora
Back to Models

Piper: The TTS Model That Doesn't Need a GPU to Exist

September 3, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Almost every voice cloning model assumes you have a GPU, VRAM to spare, and probably an internet connection to a hosted API as a fallback. Piper assumes none of that. It's built to run entirely on CPU, in real time, on hardware as modest as a Raspberry Pi — no cloud round-trip, no CUDA dependency, no waiting. If your project's actual constraint is "this needs to run on a small board with no internet connection," Piper is one of the few open-source TTS options built specifically for that reality rather than adapted to it after the fact.

What Is Piper?

Piper is an open-source neural text-to-speech system created by developer Michael Hansen under the Rhasspy voice-assistant project, with its first commit dating back to November 2022. Each Piper voice is a VITS-based model exported to the ONNX runtime, paired with a small JSON configuration file — espeak-ng handles converting input text into phonemes before the neural network generates the actual audio. That combination of a compact, ONNX-optimized model and a lightweight phonemizer is what makes Piper both natural-sounding and genuinely fast on hardware with no dedicated graphics processing at all.

Why It's So Light

A typical Piper voice model runs at roughly 15 million parameters — a fraction of the size of most cloning-focused models covered elsewhere — and a medium-quality voice packages down to a single ONNX file in the tens of megabytes. Voices ship in four quality tiers: x_low and low (16kHz output) for the smallest footprint and fastest generation, and medium and high (22.05kHz) when audio quality matters more than raw speed. That tiered design lets you deliberately trade quality for footprint depending on your specific hardware constraints, rather than being stuck with one fixed model size regardless of what you're deploying to.

The speed numbers back up the "ultra-lightweight" positioning directly: Piper synthesizes in real time on a Raspberry Pi 4 or 5 using CPU alone, and runs roughly an order of magnitude faster than real time on a modern desktop CPU — genuinely fast enough that GPU acceleration isn't just optional, it's not part of the design at all.

Language and Voice Coverage

Piper ships with more than 100 pretrained voices spanning over 35 languages, including English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Russian, Chinese, Arabic, Turkish, and Ukrainian, among others, all hosted openly on Hugging Face. Voice files follow a consistent naming convention (<language>_<region>-<name>-<quality>), making it straightforward to identify and swap between specific languages, regional accents, and quality tiers without digging through documentation each time.

Where Piper Actually Gets Used: Home Assistant and Beyond

Piper's practical footprint is concentrated in exactly the use cases its lightweight design targets: home automation, offline voice assistants, accessibility tools, and embedded kiosks where latency, privacy, and zero ongoing hosting cost matter more than studio-level vocal polish. It's a native, officially supported text-to-speech option in Home Assistant, communicating over the Wyoming protocol and auto-discovered once installed, with en_US-lessac-medium as the default voice — a strong signal of how thoroughly this project has been adopted specifically for local, privacy-conscious smart home setups rather than content production.

A Maintenance Detail Worth Knowing

Piper's original repository was archived in October 2025, and active development has since moved to a fork maintained by the Open Home Foundation's Voice team. This matters for two practical reasons: first, if you're following older tutorials or documentation, confirm whether they're pointing at the original archived repository or the actively maintained fork. Second, licensing changed along the way — the original project was MIT-licensed, while the current development fork is released under GPL-3.0, a meaningfully different set of terms if commercial use is part of your plan. Check which repository and license actually applies to the version you're deploying before building a product around it.

Getting Started with Piper

  1. Install via pip inside a virtual environment. python3 -m venv .venv && source .venv/bin/activate && pip install piper-tts avoids conflicts with system-managed Python installations, which is a common source of installation headaches on Raspberry Pi OS specifically.
  2. Download a voice model and its matching config. Each voice requires both the .onnx model file and a corresponding .json config file from the Hugging Face voice repository — grab both, and make sure they're paired correctly.
  3. Confirm which repository your instructions reference. Given the 2025 fork, verify whether a given guide is written for the original rhasspy/piper project or the current OHF-Voice/piper1-gpl fork, since setup details can differ slightly between them.
  4. Run a quick synthesis test first. Piper's CLI can generate a WAV file from a text string in one command, which is the fastest way to confirm your model and config files are correctly paired before wiring it into a larger project.
  5. Use the official Home Assistant add-on if that's your target platform. Rather than manually integrating Piper into a smart home setup, the official add-on handles discovery and configuration automatically via the Wyoming protocol.
  6. Match quality tier to your hardware, not just your ambitions. On genuinely constrained devices, x_low or low-quality voices keep generation fast and resource use minimal; reserve medium or high tiers for hardware with a bit more headroom.

Tips for Better Results

  • Start with the pip installation path unless you have a specific reason to build from source. It's the faster, more reliable route on Raspberry Pi OS, Ubuntu, and WSL on Windows, based on wide community testing across those platforms.
  • Don't default to the highest quality tier automatically. The x_low/low tiers exist specifically because many embedded use cases don't need 22.05kHz output — test the lower tiers first if you're deploying to genuinely resource-constrained hardware.
  • Treat Piper as the tool for the deployment constraint, not the content constraint. If your priority is expressive, emotionally rich narration, a larger cloning-focused model will serve you better — Piper's entire value proposition is reliability and footprint on hardware that can't run those larger models at all.
  • Check the license of your target repository before commercial deployment. Given the split between the original MIT-licensed project and the current GPL-3.0 fork, this is a genuinely important detail to verify rather than assume, especially for any commercial product.
  • Use the Wyoming protocol integration path for smart home projects specifically. If you're building anything adjacent to Home Assistant or similar voice-assistant ecosystems, working through the established Wyoming-based integration is more maintainable than a custom implementation.

Piper vs. Other Lightweight and CPU-Friendly Models

PiperKokoro-82MTypical GPU cloning models
Parameters~15M82M350M–1.7B+
Runs on CPU onlyYes, by designYes, nativelyOptional at best, GPU strongly preferred
Real-time on Raspberry PiYesNot a primary design targetNo
Voice cloningNo, fixed voice rosterNo, fixed voice rosterYes, typically the core feature
Languages35+8Varies, often fewer with this much breadth
Best fitEmbedded devices, smart home, offline kiosksCost-sensitive server-side narrationExpressive cloning, dialogue, production content

Piper's specific niche is the smallest, most resource-constrained end of the deployment spectrum — if Kokoro is built for "runs cheaply on a CPU server," Piper is built for "runs on a device that might not even have a fan."

Frequently Asked Questions

Does Piper really need no GPU at all?

Correct. Piper is CPU-only by design and uses no VRAM — it's built specifically to run offline on embedded hardware like a Raspberry Pi in real time, which is the exact use case it targets.

Can Piper clone a specific person's voice?

No. Piper works from a fixed roster of over 100 pretrained voices across 35+ languages rather than zero-shot voice cloning from a reference clip — if cloning is your core requirement, a different model is the better fit.

Is Piper still actively maintained?

Yes, though development moved from the original rhasspy/piper repository (archived in October 2025) to a fork maintained by the Open Home Foundation's Voice team, released under a different license (GPL-3.0 rather than the original's MIT).

What hardware does Piper actually need?

It runs in real time on a Raspberry Pi 4 or 5 using CPU alone, and roughly 10x faster than real time on a typical modern desktop CPU — no GPU or CUDA setup required at any point.

Is Piper free for commercial use?

It depends on which repository and version you're using — the original project was MIT-licensed, but the actively maintained fork is under GPL-3.0, which carries different obligations. Check the specific license of whichever version you deploy before commercial use.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →