EchoraEchora
Back to Models

Zonos: Open Weights, Aimed Directly at Closed Commercial Quality

September 5, 2026
•
7 min read

Zyphra didn't position Zonos as "a good open-source alternative" — the team's own claim is that it delivers expressiveness and quality on par with, or exceeding, leading commercial TTS providers, while releasing the actual weights under a fully permissive license. That's a specific, testable claim rather than a vague quality assertion, and it's backed by a genuinely large training run: more than 200,000 hours of varied speech, deliberately mixing neutral narration with highly expressive material rather than training primarily on one register.

What Is Zonos?

Zonos is an open-weight text-to-speech model released by Zyphra, a Palo Alto-based AI startup, in a February 2025 beta. It ships in two backbone variants — a standard transformer and an SSM-hybrid architecture — both at 1.6 billion parameters, and both released under the Apache 2.0 license, meaning the weights themselves, not just the code, are free for commercial use without a separate agreement.

Under the Hood: eSpeak, DAC Tokens, and Two Backbone Choices

Zonos follows a comparatively straightforward pipeline: text is normalized and converted to phonemes through eSpeak, and the resulting sequence drives token prediction over Descript Audio Codec (DAC) representations, generated through either the transformer or the SSM-hybrid backbone depending on which checkpoint you load. The practical difference between the two isn't cosmetic — the hybrid variant incorporates state-space model components alongside standard attention, which is the kind of architectural choice that typically trades off differently on speed and sequence-length handling compared to a pure transformer, so testing both on your specific workload is worth doing rather than assuming one is a strict upgrade over the other.

Voice Cloning: Speaker Embeddings vs. Audio Prefixes

Zonos supports zero-shot voice cloning from a reference clip in the 5 to 30 second range, but it's worth understanding the two distinct ways you can condition generation on a voice. A speaker embedding captures general vocal identity from the reference clip. An audio prefix — literally prepending a chunk of reference audio directly, rather than reducing it to an embedding first — can additionally elicit specific behaviors, like whispering, that are genuinely difficult to reproduce from a speaker embedding alone. If your project needs a specific delivery style that a plain reference clip isn't capturing, the audio prefix path is the more direct lever to pull.

For situations where you don't have a reference voice at all, Zonos also ships preset voice options — American and British male and female voices, along with a random-voice setting — so zero-shot cloning isn't strictly required to get usable output.

Fine-Grained Emotion and Delivery Control

Beyond voice identity, Zonos exposes direct control over speaking rate, pitch variation, maximum frequency, overall audio quality, and specific emotions — happiness, anger, sadness, and fear are explicitly supported. This level of granular, independently adjustable control across both delivery mechanics and emotional tone is a meaningful part of what the "expressive" positioning is actually built on, rather than a single blunt style parameter applied to the whole output.

Multilingual Support and Native 44kHz Output

Zonos-v0.1 covers English, Japanese, Chinese, French, and German, and — notably compared to many models in this space that output at 24kHz — generates audio natively at 44kHz, a meaningfully higher fidelity ceiling that matters if your output is going into a production pipeline where audio quality gets scrutinized rather than just needing to be intelligible.

Speed

On an RTX 4090, Zonos runs at approximately a 2x real-time factor — generating roughly 2 seconds of audio for every 1 second of compute time — putting it comfortably in usable territory for near-real-time applications on capable consumer hardware, even though it isn't purpose-built for the sub-200ms streaming latency that dedicated real-time models target.

Getting Started with Zonos

  1. Install the eSpeak dependency first. Since Zonos relies on eSpeak for text normalization and phonemization, make sure it's installed on your system before attempting to run inference — a common first-run failure point if skipped.
  2. Clone the repository and choose your backbone. git clone the Zyphra/Zonos project from GitHub, then load either Zyphra/Zonos-v0.1-transformer or Zyphra/Zonos-v0.1-hybrid depending on which architecture you want to test.
  3. Build a speaker embedding from a reference clip. A few lines of Python — loading the model, calling make_speaker_embedding() on your reference audio, then make_cond_dict() with your target text and language — cover the basic zero-shot cloning workflow.
  4. Use the Gradio interface for repeated testing. Running python gradio_interface.py avoids the overhead of reloading the model on every single generation, which the minimal Python example does by default — genuinely useful once you're iterating rather than running one sample.
  5. Try the hosted playground before committing to local setup. Zyphra maintains a hosted version at playground.zyphra.com/audio, which is a fast way to evaluate output quality and available controls before investing time in local installation.
  6. Reach for an audio prefix specifically for hard-to-clone behaviors. If a speaker embedding alone isn't producing the whispering, breathiness, or other specific delivery style you need, switch to the audio-prefix conditioning method instead.

Tips for Better Results

  • Test both backbone variants on your actual use case. The transformer and SSM-hybrid models aren't simply "old vs. new" — they represent a genuine architectural trade-off, so benchmarking both on your specific content type is more reliable than assuming one is universally better.
  • Use a genuinely clean reference clip within the 5–30 second range. Since this window is what zero-shot cloning is tuned around, a well-recorded sample at a length within that range will outperform either a much shorter or unnecessarily padded-out clip.
  • Combine emotion and delivery controls deliberately, not just one at a time. Since speaking rate, pitch, and emotion are independently adjustable, layering them thoughtfully (a faster pace paired with a specific emotional tone, for instance) is where Zonos's expressiveness claim actually shows up, rather than adjusting a single parameter in isolation.
  • Use preset voices when you don't need cloning specifically. If reproducing a particular real voice isn't the goal, the built-in American/British male and female presets save the step of sourcing and preparing a reference clip.
  • Watch for Zyphra's newer ZONOS2 release if you need broader language coverage or lower latency. Zyphra has since released a follow-up model trained on a substantially larger dataset with expanded language support and a mixture-of-experts architecture aimed at lower latency — worth checking if your specific requirements outgrow what v0.1 covers.

Zonos vs. Other Expressive Cloning Models

ZonosChatterboxF5-TTS
Backing organizationZyphraResemble AIIndependent research (SJTU/Cambridge/Geely)
Parameters1.6B (two backbone variants)~0.5BNot primarily framed by parameter count
Reference clip for cloning5–30 seconds5–10 seconds~5–15 seconds
Emotion controlExplicit: happiness, anger, sadness, fearExaggeration parameterNot a dedicated feature
Native output sample rate44kHzNot specified at this fidelityStandard TTS rates
LanguagesEnglish, Japanese, Chinese, French, GermanEnglish (Multilingual: 23)Primarily English-focused
LicenseApache 2.0MITCC-BY-NC (non-commercial)

Zonos's specific edge is combining a genuinely large, dual-register training corpus with fully open, commercially permissive weights at 44kHz native fidelity — a combination that's harder to find bundled together elsewhere, even if its five-language coverage is narrower than some multilingual-focused competitors.

Frequently Asked Questions

Is Zonos actually free for commercial use?

Yes. Both the transformer and SSM-hybrid model weights are released under the Apache 2.0 license, which permits commercial use without a separate agreement.

What's the difference between the transformer and hybrid Zonos models?

The hybrid model incorporates state-space model components alongside standard transformer attention, representing a different architectural trade-off rather than a straightforward upgrade — testing both on your specific use case is the reliable way to pick.

How much reference audio does Zonos need for voice cloning?

Roughly 5 to 30 seconds, using either a speaker embedding (for general vocal identity) or an audio prefix (better suited for specific behaviors like whispering).

What languages does Zonos support?

English, Japanese, Chinese, French, and German in the v0.1 release.

Is there a hosted way to try Zonos without installing anything?

Yes. Zyphra maintains a hosted playground at playground.zyphra.com/audio where you can test the model's voice cloning and expressive controls directly.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.

Create a Voice Clone →