Parler-TTS: Describe the Voice You Want, Instead of Recording It
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Every cloning model asks for the same thing first: a reference audio clip. Parler-TTS, developed by Hugging Face, skips that requirement entirely — instead of a voice sample, you write a sentence describing the voice you want, and the model generates speech matching that description. "A female speaker with a slightly low-pitched voice, speaking quickly in a clear, close-up recording" is a complete voice specification here, no microphone required.
What Is Parler-TTS?
Parler-TTS is an open-source text-to-speech project from Hugging Face, built as a reproduction of research on natural-language guidance for high-fidelity TTS, originally from Stability AI and the University of Edinburgh. What sets the project apart from most open-source releases in this space is how completely open it actually is: the datasets, preprocessing pipeline, training code, and model weights are all released publicly under a permissive license, rather than just the final checkpoint. That's a meaningfully more transparent release than most TTS projects offer, and it's specifically what lets the community retrain, fine-tune, or fully reproduce the project's results rather than just running inference on a black box.
How Natural-Language Voice Control Actually Works
Parler-TTS takes two separate text inputs rather than one — the text you want spoken, and a plain-language description of how it should sound. That description can control speaker gender, pitch, speaking rate, and recording characteristics like background noise or reverberation, all through ordinary sentences rather than numeric sliders or reference audio. A description like "a male speaker delivers a monotone speech at a fast pace in a noisy recording" is parsed and applied directly to the generated output.
Making that work required solving a real data problem: audio datasets don't come with natural-language descriptions attached. Hugging Face built a companion project, DataSpeech, specifically to bridge that gap — it tags speech samples with objective feature labels (noise level, speaking rate, reverberation, and similar attributes), then feeds those tags to a large language model to generate natural-sounding descriptions from them. That LLM-generated description, paired with its corresponding audio, becomes training data teaching Parler-TTS to associate descriptive language with actual acoustic characteristics.
What You Can Actually Control
The base models support description-driven control over gender, pitch, speaking rate, and acoustic environment (background noise, reverberation, recording clarity) — attributes that are largely universal across any speaker. Emotional expressiveness is a further, fine-tuning-based capability: Hugging Face released a version fine-tuned on the Expresso dataset — a high-quality expressive speech corpus featuring four speakers — specifically to teach the model to follow emotion and performance-style prompts, extending the same natural-language control mechanism into more expressive territory than the base models cover by default.
The Model Family
Parler-TTS has been released in a few distinct sizes and training scales, each representing a step up in either data or capability:
- Parler-TTS Mini v0.1 — the first release, trained on roughly 10,500 hours of audio, establishing the core natural-language description approach.
- Parler-TTS Mini v1 / v1.1 — trained on a larger 45,000-hour dataset, with v1.1 specifically improving the prompt tokenizer (based on the Llama 2 tokenizer, with a larger vocabulary and byte-fallback handling) to simplify future multilingual training, even though the underlying model and training data are otherwise identical to v1.
- Parler-TTS Large v1 — a 2.2-billion-parameter model, also trained on the 45,000-hour dataset, offering the highest quality tier in the family at a correspondingly larger compute footprint.
Getting Started with Parler-TTS
- Install the library.
pip install git+https://github.com/huggingface/parler-tts.gitgets you the inference package directly from the source repository. - Load your chosen model and tokenizer. Use
ParlerTTSForConditionalGeneration.from_pretrained()with the specific checkpoint you want (Mini v0.1, Mini v1.1, or Large v1), paired with the matchingAutoTokenizer. - Write two separate prompts. One variable holds the text you want spoken; a second holds your natural-language description of the voice and delivery — both get tokenized separately before being passed into generation.
- Use two tokenizers if you're on v1.1 or later. Unlike earlier versions, Mini v1.1 and Large v1 use separate tokenizers for the description and the spoken text respectively — check which version's documentation you're following, since mixing up the single-tokenizer approach from v0.1 with a newer checkpoint will cause issues.
- Fine-tune on your own annotated dataset if you need a specific capability. The DataSpeech pipeline is designed specifically to help you generate the natural-language descriptions needed to fine-tune Parler-TTS toward a new voice style, language, or expressive capability, following the same approach used to create the Expresso-fine-tuned version.
Tips for Better Results
- Be specific rather than vague in your voice description. "A female speaker with a slightly low-pitched voice delivers her words in a very clear, close-up, and slightly fast recording" gives the model considerably more to work with than "a normal voice," and specificity across each controllable dimension (gender, pitch, rate, recording quality) tends to produce more consistent results.
- Avoid describing nationality in your prompts. Community fine-tuning work has found that including nationality in training data or inference prompts can actually degrade output quality rather than improve accent accuracy — omit it and let other descriptive terms (accent, tone, pacing) carry that information instead.
- Reach for the Expresso fine-tune specifically for emotional or performance-style prompts. The base models are tuned for the more universal attributes (gender, pitch, rate, environment); if your project needs a specific emotional delivery, the Expresso-based checkpoint is built for that rather than the general-purpose base model.
- Match model size to your quality needs. Mini v1.1 is the more practical default for most projects; reach for Large v1's 2.2B parameters specifically when maximum quality matters more than inference speed and resource use.
- Consider training your own fine-tune for non-English languages. Since Parler-TTS's base models are primarily English-focused, community efforts (a French fine-tune is one documented example) have shown the DataSpeech-plus-fine-tuning approach can extend the project to other languages, though open dataset quality and quantity remain the main practical bottleneck.
Parler-TTS vs. Reference-Audio Cloning Models
| Parler-TTS | XTTS-v2 | F5-TTS | |
|---|---|---|---|
| Voice control method | Natural-language text description | Reference audio clip (~6 seconds) | Reference audio clip (~5–15 seconds) |
| No audio sample needed | Yes | No | No |
| Clone a specific real person's voice | Not the design goal | Yes, that's the core purpose | Yes, that's the core purpose |
| Full pipeline openness (data + code + weights) | Yes | Code and weights only | Code and weights only |
| Model sizes | Mini (10.5K/45K hrs) to Large (2.2B params) | Single current version | Single current version |
| License | Apache 2.0 | Coqui Public Model License (non-commercial) | CC-BY-NC (non-commercial) |
Parler-TTS's specific niche is generating a plausible, well-specified voice on demand without needing a real recording to clone — useful when you want control over a voice's characteristics rather than reproduction of one specific person's identity, and its fully open training pipeline is a genuine rarity among projects at this scale.
Frequently Asked Questions
Do I need a reference audio clip to use Parler-TTS?
No. Unlike voice-cloning models, Parler-TTS generates speech based on a natural-language text description of the desired voice and delivery style, with no audio sample required.
What voice characteristics can I actually control?
The base models support gender, pitch, speaking rate, and acoustic environment factors like background noise and reverberation, all specified through plain-language descriptions. Emotional expressiveness is available through the Expresso fine-tuned version specifically.
Is Parler-TTS free for commercial use?
Yes. It's released under the Apache 2.0 license, one of the more permissive options among open-source TTS models, covering the code, training pipeline, and weights alike.
What's the difference between Mini v1 and Mini v1.1?
They're trained on identical data with the same configuration; v1.1's only change is an improved prompt tokenizer with a larger vocabulary and byte-fallback support, intended to simplify future multilingual training.
Can Parler-TTS clone a specific real person's voice?
Not as its core design. Parler-TTS is built around generating a voice matching a text description rather than reproducing a specific reference speaker's identity — for cloning a real person's voice, a dedicated cloning model is the better fit.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.