EchoraEchora
Back to Models

MetaVoice: Built to Get the Emotional Rhythm Right, Not Just the Words

September 7, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Plenty of TTS models pronounce every word correctly and still sound like they're reading a shopping list. MetaVoice, from the team at metavoice.io, was built around a narrower, more specific priority: emotional speech rhythm and tone in English, with an explicit design goal its own documentation states plainly — no hallucinations. Rather than chasing the widest possible language coverage, it concentrates on making English delivery sound genuinely felt rather than merely correct.

What Is MetaVoice?

MetaVoice is a 1.2-billion-parameter open-source text-to-speech model trained on 100,000 hours of speech data, released under the fully permissive Apache 2.0 license with no usage restrictions. Under the hood, it predicts EnCodec audio tokens from text and speaker information through a multi-step pipeline: a causal GPT-style component, a non-causal transformer, and multi-band diffusion working together to reconstruct the final waveform, with KV-caching and batching built in to keep inference practical rather than purely research-grade.

Zero-Shot Cloning for American and British English

MetaVoice's most immediately usable capability is zero-shot voice cloning specifically tuned for American and British English accents, requiring just 30 seconds of reference audio to reproduce a target speaker's voice with no fine-tuning step required. This narrower accent focus — rather than a broad multilingual voice bank — is a deliberate trade-off: depth on two well-represented English accents instead of shallow coverage spread across many languages.

Cross-Lingual Voice Cloning Through Fine-Tuning

Beyond its zero-shot English capability, MetaVoice also supports cross-lingual voice cloning, though this path requires fine-tuning rather than working out of the box. The project's own documentation reports success with as little as one minute of training data when adapting the model to Indian-accented speakers, which is a genuinely small amount of data for a cross-lingual adaptation task. It's worth being precise about what this means in practice: the base model itself isn't natively multilingual in the way a model trained from the ground up on many languages would be — extending it beyond its core American and British English focus is an adaptation you perform yourself, with your own reference data, rather than a capability that ships ready-to-use for arbitrary languages.

Speed and Deployment Options

MetaVoice supports long-form synthesis alongside its shorter-clip cloning use cases, and offers experimental quantization modes (int4 and int8) for faster inference at some cost to audio quality — int4 runs roughly twice as fast as the standard bf16/fp16 precision, making it the more practical choice if you specifically need the quantized speed boost, since int8 has been reported as actually slower than full precision for reasons the project hasn't fully resolved. On modern GPU architectures (Ampere, Ada Lovelace, and Hopper), once the model has been compiled for inference — a one-time startup cost of roughly 30 to 90 seconds depending on your hardware — generation runs faster than real time, with a real-time factor below 1.0.

Getting Started with MetaVoice

Setup involves cloning the project's repository and following its documented installation steps, which include options for running the interactive Python inference session directly, deploying via Docker Compose with a bundled API server, or standing up the included web UI for a browser-based workflow. The project can be deployed on any major cloud provider using its inference server, making self-hosting straightforward whether you're testing locally or running at production scale. Given the torch.compile warmup step, plan for that initial startup delay on first run rather than expecting instant generation from a cold start.

Tips for Better Results

  • Use a genuinely clean 30-second reference clip for zero-shot cloning. Since this is the core supported workflow for American and British accents, a well-recorded sample at roughly that length will outperform either a shorter or noisier clip.
  • Treat cross-lingual adaptation as a fine-tuning project, not a flag you flip. If you need a voice or accent outside the base American/British English focus, budget time for gathering your own small reference dataset and following the documented fine-tuning process, rather than expecting it to work zero-shot.
  • Reach for int4 quantization specifically when speed matters more than peak quality. Given its roughly 2x speed advantage over standard precision, it's the more sensible quantized option to test first if inference speed is your bottleneck.
  • Budget for the initial compilation delay in production planning. Since the model is compiled on startup for faster-than-real-time inference afterward, factor that one-time delay into any cold-start latency expectations for your deployment.

MetaVoice vs. Other Accent-Focused Cloning Models

MetaVoiceXTTS-v2Chatterbox
Core language focusEnglish (American, British)17 languagesEnglish (Multilingual variant: 23)
Zero-shot cloningYes, ~30 second referenceYes, ~6 second referenceYes, 5–10 second reference
Cross-lingual/accent adaptationVia fine-tuning (e.g., Indian English)Native, built into the base modelNative, via Multilingual variant
Emotional expressivenessCore design priorityAutomatic via reference encoderExplicit exaggeration parameter
Parameters1.2BNot publicly emphasized~0.5B
LicenseApache 2.0Coqui Public Model License (non-commercial)MIT

MetaVoice's specific niche is depth over breadth — rather than competing on language count, it concentrates on making English speech emotionally convincing, with cross-lingual reach available as a fine-tuning path for teams willing to do that extra work themselves.

Frequently Asked Questions

Does MetaVoice support Chinese or Chinese-English code-switching?

Not out of the box. Its documented focus is American and British English, with cross-lingual capability available only through fine-tuning on your own data — there's no documented native support for Chinese or mixed-language code-switching in the base model.

How much reference audio does MetaVoice need for cloning?

About 30 seconds for zero-shot cloning of American or British English voices. Cross-lingual or accent adaptation through fine-tuning has been reported working with as little as one minute of data.

Is MetaVoice free for commercial use?

Yes. It's released under the Apache 2.0 license with no usage restrictions, one of the more permissive options among open-source TTS models.

What does "no hallucinations" mean as a design goal here?

It reflects a specific engineering priority to avoid the model generating unintended words, sounds, or artifacts not present in the input text — a reliability-focused goal alongside its emotional-expressiveness priority, rather than a claim about eliminating every possible generation error.

Does quantization affect audio quality?

Yes. Both the int4 and int8 quantization modes trade some audio quality for faster inference, with int4 offering roughly a 2x speed improvement over standard precision.

Create a Similar Voice from Reference Audio

For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.

Create a Voice Clone →