Bark: The Model That Treats Speech as Just One Kind of Sound
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Ask most TTS models for a laugh and you'll get either silence or something that sounds like a sound-effect library was crudely spliced in. Bark, the open-source model from Suno, handles it differently because it was never built to be a narrow text-to-speech engine in the first place — it's a generative text-to-audio model, and speech is just one of the things it happens to be good at. Drop [laughs] into a sentence and Bark performs the laugh itself, in context, the same way it generates the words around it.
What Is Bark?
Bark is a Transformer-based generative audio model created by Suno, released under the MIT license. Rather than following the traditional TTS pipeline — text to phonemes to acoustic features to waveform — Bark works more like a language model that happens to output audio tokens instead of text tokens. It converts written prompts directly into a 24kHz waveform through a three-stage autoregressive pipeline: first generating semantic tokens representing the intended content, then coarse acoustic tokens, then fine acoustic tokens, using Meta's EnCodec for the underlying audio tokenization. There's no phoneme conversion step at any point — text goes in, and increasingly detailed audio representations come out, end to end.
Bracket Tags: How Laughter and Sighs Actually Work
This is Bark's defining feature, and it's genuinely simple to use. Non-verbal sounds and emotional cues are triggered by inserting tags directly into the input text: [laughter], [sighs], [gasps], and similar markers prompt the model to generate that sound at that exact point in the sentence, performed by whichever voice is currently speaking rather than dropped in as a separate audio clip. Beyond bracket tags, Bark also reads plain formatting as delivery cues — CAPITALIZING a word pushes the model toward emphasizing it, and trailing off a sentence with "..." produces a natural hesitation in the delivery, rather than requiring a dedicated parameter for either effect.
More Than Speech
Because Bark generates audio as a general category rather than speech specifically, it can also produce simple music, background ambience, and basic sound effects directly from a text prompt — capabilities that fall outside what a dedicated TTS model is built to do at all. It supports more than a dozen languages with natural code-switching, meaning a script that shifts between languages mid-sentence doesn't require separate generation calls, and it ships with a range of built-in speaker presets covering different vocal styles rather than a single default voice.
What Bark Doesn't Do (Read This Before You Start)
A few honest limitations are worth knowing upfront, because they shape what Bark is actually good for:
It's not primarily a voice-cloning model. Unlike most cloning-focused systems, Bark's standard workflow is built around its built-in speaker presets rather than reproducing an arbitrary reference voice from a short clip — if voice cloning of a specific person is your core requirement, a model purpose-built for that job will serve you better.
Output can genuinely surprise you. Because Bark is a fully generative model rather than a constrained TTS system, it can occasionally deviate from the input prompt in unexpected ways — skipping a phrase, adding an unrequested sound, or otherwise not saying exactly what you typed. This is a structural property of the architecture, not a bug you can configure away.
It's slow, and generation length is limited per call. Bark's autoregressive pipeline means a single sentence commonly takes 5 to 15 seconds to generate on GPU, and each generation call is limited to roughly 13-14 seconds of audio — longer content needs to be chunked into multiple calls rather than generated in one pass.
Getting Started with Bark
- Install via pip or clone the repository.
pip install git+https://github.com/suno-ai/bark.gitgets you running quickly, or clone the repo directly if you want to inspect or modify the code. - Check your PyTorch and CUDA versions. Bark needs PyTorch 2.0 or newer, with CUDA 11.7 or 12.0 support if you're running on GPU.
- Use the small model variant on limited VRAM. Setting the environment variable
SUNO_USE_SMALL_MODELS=Truebefore loading the model lets Bark run on roughly 8GB of VRAM, at some cost to output quality. - Write your prompt with tags inline. A basic generation is a few lines: load the model, pass a text string containing your desired bracket tags, and write out the resulting waveform — no separate configuration step needed for the non-verbal cues themselves.
- Chunk longer scripts into multiple generations. Given the roughly 13-14 second limit per call, split longer narration into sentence- or paragraph-sized pieces and concatenate the resulting audio rather than expecting one call to handle a full paragraph.
- Try the Hugging Face Transformers integration if you'd rather skip the standalone repo. Bark has been supported in Transformers since version 4.31.0, which can simplify integration if your project already depends on that library.
Tips for Better Results
- Use tags sparingly and in context. A
[laughs]tag lands naturally right after something genuinely funny in the script; scattering non-verbal tags throughout dilutes the effect and can make delivery feel scripted rather than reactive. - Test a preset speaker before assuming you need a custom voice. Since Bark isn't built around zero-shot cloning, browsing its built-in speaker presets first is usually faster than trying to force an arbitrary reference voice into the pipeline.
- Regenerate rather than fight unpredictable output. Given Bark's generative nature, an occasional odd result is often faster to fix by simply regenerating the same prompt than by trying to prevent it through prompt engineering alone.
- Plan for the per-call length limit from the start. Structuring your script into natural chunks upfront saves rework compared to writing one long passage and discovering the length ceiling partway through a project.
- Use it specifically where expressiveness matters more than precision. Character voice acting, animated content, and creative audio projects are where Bark's non-verbal range genuinely shines — for precise, predictable narration, a more conventional TTS model may be the better fit.
Bark vs. Other Expressive TTS Models
| Bark | Chatterbox | CosyVoice3 | |
|---|---|---|---|
| Core purpose | Generative text-to-audio | Voice cloning with emotion control | Multilingual cloning with dialects |
| Non-verbal sounds | Native bracket tags, plus music/effects | Paralinguistic tags (Turbo variant) | Not a dedicated feature |
| Voice cloning | Not the primary workflow | Yes, zero-shot | Yes, zero-shot |
| Output predictability | Can deviate from prompt | More consistent | More consistent |
| Languages | 13+ | English (Multilingual: 23) | 9 + 18 Chinese dialects |
| License | MIT | MIT | Open, hosted API available |
Bark's specific niche is expressive, character-driven audio where laughter, sighs, and tonal shifts matter more than reproducing one exact person's voice — for that particular job, few open-source models generate non-verbal sound as naturally as it does.
Frequently Asked Questions
How do I make Bark generate laughter or sighs?
Insert bracket tags like [laughter] or [sighs] directly into your input text at the point where you want the sound to occur — Bark performs it in the current speaker's voice rather than requiring a separate audio effect.
Can Bark clone a specific person's voice?
Not as its primary design. Bark is built around a set of preset voices rather than zero-shot cloning from a reference clip — for cloning a specific real voice, a dedicated cloning model is a better fit.
Is Bark free for commercial use?
Yes. Bark is released under the MIT license, which permits commercial use without royalties or separate licensing agreements.
Why does Bark sometimes generate something different from what I typed?
Bark is a fully generative audio model rather than a constrained TTS system, so it can occasionally deviate from the exact prompt — this is a known characteristic of its architecture rather than a configuration issue.
How long can a single Bark generation be?
Each generation call is limited to roughly 13-14 seconds of audio; longer content needs to be split into multiple chunks and generated separately.
Create Speech Online with Echora
For a browser-based workflow with a similar goal, use Echora’s Text to Speech. Enter up to 5,000 characters, choose a built-in or saved custom voice, add supported delivery tags when needed, and adjust voice stability. You can then preview the result and download the finished MP3.