EchoraEchora
Back to Models

Fish Speech: The TTS Project That Gives You Two Roads to a Better Voice

August 29, 2026
•
7 min read

Turn text and recordings into audio you can hear

Dubbing, stems, covers, and sound effects — all in the browser.

  • Text to speech, plus voice cloning
  • Split stems, or generate a cover
  • Sign up to preview — listen first, decide later
Start creating for free

Most cloning models offer exactly one path to a custom voice: drop in a reference clip and hope the zero-shot result is good enough. Fish Speech, the open-source text-to-speech project from Fish Audio, gives you a second option — fine-tune the model directly on a voice when zero-shot isn't precise enough — alongside a level of emotional and stylistic control that goes well beyond flat narration.

What Is Fish Speech?

Fish Speech is built around a DualAR architecture — a dual autoregressive transformer design — and trained on well over a million hours of multilingual audio. The project's 1.5 release covers 13 languages, including English, Chinese, and Japanese, and posted strong independent benchmark results: a TTS Arena ELO score of 1339 (ranking first among open-source models at the time), a 3.5% word error rate and 1.2% character error rate in English, and a 1.3% character error rate in Chinese.

The project has continued to evolve past 1.5. Fish Audio Text to Speech has since released newer generations under its OpenAudio series, with Fish Audio S2 now standing as the flagship model — trained on more than 10 million hours of audio and currently holding top rankings on evaluations like EmergentTTS-Eval, Seed-TTS Eval, and the Audio Turing Test. If you're specifically searching for Fish Speech 1.5, it remains a genuinely solid, lighter-weight, self-hostable option; if raw quality and fine-grained emotional control are the priority, S2 is the newer generation built for that.

Zero-Shot Cloning and Fine-Tuning, Not Just One or the Other

This is Fish Speech's defining trait relative to many cloning-focused competitors. Zero-shot cloning works the way you'd expect — a 10 to 30 second reference clip is enough to capture a speaker's timbre, style, and general emotional tendency, with no training step required. But Fish AI Text to Speech also supports genuine fine-tuning on top of that: when a voice needs to be more precise or consistent than zero-shot cloning can reliably deliver — a recurring brand voice, a specific character across a long project — you can fine-tune the base model directly on your own audio data rather than settling for whatever the reference clip produces.

This dual approach matters in practice: zero-shot is fast and good enough for most one-off needs, but fine-tuning is what production teams reach for when a voice has to be dependable across hundreds of generations rather than a single clip.

Genuinely Fine-Grained Emotional Control

Where many TTS models offer a single "exaggeration" slider at best, Fish Audio S2 supports sub-word-level control over prosody and emotion through natural-language tags dropped directly into the text — things like [whisper], [excited], or [angry] — letting a single sentence shift tone mid-delivery rather than applying one emotional setting to the whole clip. The model also natively handles multi-speaker and multi-turn conversation generation, and evaluates its own output against multiple reward signals — semantic accuracy, instruction adherence, acoustic quality, and timbre similarity — during training, which is part of why its emotional delivery holds up rather than sounding like a single tag was bolted onto otherwise flat narration.

Because the underlying Dual-AR architecture is structurally similar to a standard language model, S2 also natively supports high-throughput LLM inference techniques like continuous batching, paged KV cache, and prefix caching — a meaningful advantage if you're serving generation requests at scale rather than running one-off clips.

Getting Started with Fish Speech

  1. Pick your version based on your priority. Fish Speech 1.5 is the lighter, well-documented option for self-hosting on modest hardware; Fish Audio S2 is the current flagship for maximum quality and fine-grained emotional control.
  2. Clone the repository. git clone https://github.com/fishaudio/fish-speech gets you the code, documentation, and installation instructions for whichever version you're targeting.
  3. Check your hardware against the model's needs. Fish Speech 1.5 wants roughly 12GB of VRAM minimum (24GB recommended for production workloads); confirm your GPU fits before committing to a deployment plan.
  4. Try zero-shot cloning first. Provide a 10–30 second reference clip and generate directly — no training required — before deciding whether fine-tuning is actually necessary for your use case.
  5. Fine-tune only when zero-shot isn't precise enough. If a specific voice needs to be consistent across a large volume of content, follow the project's fine-tuning documentation rather than repeatedly regenerating from a single reference clip and hoping for consistency.
  6. Use the SGLang or vLLM-Omni server for production S2 deployments. Both are documented specifically for Fish Audio S2 and take advantage of its LLM-like architecture for meaningfully better throughput than a basic inference script.

Tips for Better Results

  • Match your reference clip length to your goal. Shorter clips within the 10–30 second range are usually enough for a quick zero-shot test; if the result isn't stable enough, that's your signal to move to fine-tuning rather than trying progressively longer reference clips.
  • Use emotion tags deliberately, not everywhere. Dropping [whisper] or [excited] into every sentence dilutes the effect — reserve tags for the specific moments where a shift in delivery actually matters to the content.
  • Don't skip the license review. Fish Speech's model weights have historically been released under research-oriented, non-commercial-leaning licenses (the 1.5 weights under CC BY-NC-SA, and S2 under Fish Audio's own research license) — read the current terms carefully before any commercial deployment, since they differ meaningfully from a permissive license like MIT or Apache 2.0.
  • Reach for the mixed-language handling if your content switches naturally. Fish Speech is built to read mixed Chinese-English text fluently without extra configuration — useful for bilingual scripts rather than needing to split and stitch audio from two separate generations.
  • Benchmark both versions on your actual content before choosing. TTS Arena rankings and published error rates are a useful starting signal, but voice quality is content-dependent — test 1.5 and S2 on your specific script style before committing to a production pipeline.

Fish Speech vs. Other Cloning-Focused Models

Fish Speech (1.5 / S2)ChatterboxCosyVoice3
Zero-shot cloningYes, 10–30s referenceYes, 5–10s referenceYes, 3–10s reference
Fine-tuning supportYes, nativeNot a primary workflowNot a primary workflow
Emotion controlSub-word tags (S2)Exaggeration parameterInstruction-based
Languages13 (1.5) / broader (S2)English (Multilingual: 23)9 + 18 Chinese dialects
ArchitectureDual autoregressive transformerDiffusion-basedSupervised semantic tokens
LicenseResearch/non-commercial-leaningMITOpen, hosted API available

Fish Speech's specific niche is teams that need a voice to be more consistent than a single reference clip can guarantee — the fine-tuning path is the differentiator most single-purpose zero-shot cloning models don't offer at all.

Frequently Asked Questions

What's the difference between Fish Speech 1.5 and Fish Audio S2?

Fish Speech 1.5 is the earlier, lighter-weight release in the same project lineage, well-suited to self-hosting on modest hardware. Fish Audio S2 is the current flagship, trained on far more data and offering more fine-grained emotional control through natural-language tags, currently ranking at the top of several independent TTS evaluations.

Can Fish Speech actually be fine-tuned, or is it zero-shot cloning only?

Both are supported. Zero-shot cloning works from a 10–30 second reference clip with no training step; fine-tuning is available for cases where you need a specific voice to be more consistent than zero-shot cloning reliably delivers.

How do the emotion tags in Fish Audio S2 work?

You place natural-language tags like [whisper], [excited], or [angry] directly in the input text, and the model adjusts delivery at that specific point in the sentence rather than applying one emotional setting to the whole output.

Is Fish Speech free for commercial use?

Check the current license carefully before assuming so. Fish Speech 1.5's weights have been released under a CC BY-NC-SA license (non-commercial), and Fish Audio S2 is governed by Fish Audio's own research license — neither is as permissive as MIT or Apache 2.0, so commercial deployment terms need direct verification.

What hardware do I need to self-host Fish Speech?

Fish Speech 1.5 needs roughly 12GB of VRAM at minimum, with 24GB recommended for production workloads; check current documentation for S2's specific requirements, as its larger training scale generally implies higher hardware needs.

Create Speech Online with Echora

For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.

Create Speech from Text →