NeuTTS Air: Voice Cloning Small Enough to Fit on a Raspberry Pi
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most genuinely lightweight, CPU-friendly TTS models make a specific trade-off: small footprint in exchange for giving up voice cloning, working instead from a fixed roster of preset voices. NeuTTS Air, from Neuphonic, refuses that trade-off. It's under a billion parameters, quantized to run comfortably on a phone or a Raspberry Pi, and it still clones a target voice from just 3 seconds of reference audio — a combination Neuphonic describes as the first genuinely realistic, fully on-device TTS speech language model to pull off instant cloning at this size.
What Is NeuTTS Air?
NeuTTS Air is an open-source, Apache 2.0-licensed text-to-speech model released by Neuphonic in October 2025, built around a Qwen2-based language model backbone (listed at 748 million parameters on its official model card, commonly described as "0.5B-class") paired with NeuCodec, Neuphonic's own neural audio codec. The whole point of the design is on-device deployment — running entirely on local hardware, with no audio or text ever needing to leave the device unless you choose to send it somewhere.
Functional Advantages of NeuTTS Air
- Instant voice cloning from as little as 3 seconds of reference audio, a genuinely uncommon capability among models this size, most of which work from a fixed set of preset voices instead.
- GGUF quantization (Q4 and Q8) built specifically for edge deployment, running through
llama.cpp/llama-cpp-pythonon phones, laptops, and single-board computers like the Raspberry Pi. - NeuCodec, a highly efficient audio codec operating at just 0.8kbps using a single codebook at a 50Hz frame rate — the specific design choice that keeps output quality high despite the model's small overall footprint.
- Real-time generation on genuinely modest hardware, with published benchmarks showing over 2x real-time throughput on an AMD Ryzen 9 CPU, and a Samsung Galaxy A25 smartphone reaching close to real-time speed.
- Built-in Perth watermarking embedded automatically in every generated audio file, applied by default rather than as an optional setting.
- Fully open under Apache 2.0, permitting commercial use without a separate licensing agreement.
How the Architecture Balances Realism and Footprint
NeuTTS Air's Qwen2-based backbone handles phonemization (via an espeak-ng front end), prosody modeling, and conditioning the output on both the input text and the reference speaker's style — essentially the full linguistic and expressive reasoning work, packed into a sub-billion-parameter language model rather than a much larger one. NeuCodec then handles the acoustic side: because it compresses audio using only a single codebook rather than the multiple stacked codebooks many neural codecs rely on, it achieves notably efficient compression (around 0.8kbps at 24kHz output) without the typical quality cost of using fewer codebooks. The model's 2,048-token context window covers roughly 30 seconds of audio, including the reference prompt — plenty for short-form conversational use, though a real constraint to keep in mind if you're planning long-form narration rather than assistant-style dialogue.
Getting Started with NeuTTS Air
- Install with the GGUF extras for quantized CPU inference.
uv sync --extra ggufpulls inllama-cpp-python, which is required to run the Q4/Q8 GGUF checkpoints; adduv sync --extra onnxas well if you want the ONNX decoder option instead of or alongside GGUF. - Choose your backbone checkpoint from Hugging Face.
neuphonic/neutts-airis the full-precision model;neuphonic/neutts-air-q4-ggufandneuphonic/neutts-air-q8-ggufare the quantized versions built specifically for edge and CPU deployment. - Prepare a reference clip and its transcript. NeuTTS Air needs both a clean, mono reference WAV file (3 to 15 seconds is the documented sweet spot) and the matching transcript text — natural, continuous speech with few pauses works better than a clipped or heavily edited sample.
- Run the basic example script to confirm your setup.
python -m examples.basic_example --input_text "your text here" --ref_audio samples/dave.wav --ref_text samples/dave.txt(adjusting the--backboneflag for your chosen checkpoint) is the documented first test. - Use the Python API directly for integration.
from neuttsair.neutts import NeuTTSAir, then instantiate with your chosenbackbone_repoandcodec_repo(neuphonic/neucodecis the standard codec choice), calltts.encode_reference()on your reference clip, and pass the result intotts.infer()alongside your target text. - Check the NeuTTS-Air Hugging Face collection for additional backbone options before assuming the default checkpoint is your only choice — several variants are documented there for different deployment needs.
Tips for Better Results
- Keep your reference clip clean and conversational. A natural monologue or conversational sample with minimal pauses helps the model capture tone accurately — background music, overlapping speech, or heavily produced audio will noticeably hurt cloning quality.
- Stay within the 2,048-token context window when planning content. Since this covers roughly 30 seconds including your reference prompt, structure longer content into multiple generation calls rather than expecting one pass to handle an extended passage.
- Use Q4 GGUF for the most constrained devices, Q8 when you have a bit more headroom. The quantization level directly trades a small amount of quality for speed and memory footprint — pick based on your specific target hardware rather than defaulting to one tier.
- Consider NeuTTS Nano if even 748M parameters is too much. Neuphonic's smaller sibling model, at 229M total parameters (120M active), targets even more constrained edge deployments than NeuTTS Air itself.
NeuTTS Air vs. Other Lightweight On-Device Models
| NeuTTS Air | Piper | Kokoro-82M | |
|---|---|---|---|
| Parameters | ~748M (Qwen2-based) | ~15M | 82M |
| Voice cloning | Yes, from ~3 seconds | No, fixed voice roster | No, fixed voice roster |
| Deployment format | GGUF (Q4/Q8) via llama.cpp | ONNX | PyTorch / ONNX |
| Real-time on smartphone-class hardware | Yes (documented Galaxy A25 benchmark) | Yes | Not a primary design target |
| Built-in watermarking | Yes (Perth) | No | No |
| Languages | English (strongest), Spanish, German, French | 35+ | 8 |
| License | Apache 2.0 | MIT (original) / GPL-3.0 (current fork) | Apache 2.0 |
NeuTTS Air's specific niche within the lightweight, on-device category is being one of the few options that keeps zero-shot voice cloning intact at this size — most models this small give up cloning entirely in exchange for their small footprint, which is exactly the trade-off NeuTTS Air was built to avoid.
Frequently Asked Questions
Does NeuTTS Air need a GPU?
No. It's optimized for CPU inference through GGUF quantization and runs well on modern laptops or devices like a Raspberry Pi 4; GPU support is available through llama-cpp-python with CUDA or Metal if you want it, but it isn't required.
How much reference audio does NeuTTS Air need for voice cloning?
As little as 3 seconds, with 3 to 15 seconds of clean, mono audio documented as the practical sweet spot, along with a matching transcript of that reference clip.
Is NeuTTS Air free for commercial use?
Yes. It's released under the Apache 2.0 license, which permits commercial use without a separate licensing agreement.
Why does NeuTTS Air need espeak?
It uses espeak-ng to convert input text into phonemes as a preprocessing step — a critical part of how the language-model backbone understands what sounds it needs to generate.
Create a Similar Voice from Reference Audio with Echora
For a browser-based workflow with a similar goal, use Echora's Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.