Sesame CSM 1B: Speech That's Grounded in the Conversation, Not Just the Text
- What Is CSM 1B?
- Under the Hood: Two Transformers and the Mimi Codec
- The Real Differentiator: Grounding Speech in Actual Conversation
- An Honest Note: The Open Checkpoint vs. the Viral Demo
- What CSM Can't Do
- Getting Started with CSM 1B
- Tips for Better Results
- CSM 1B vs. Other Llama-Backbone Speech Models
- Frequently Asked Questions
- Turn a Multi-Speaker Script into Audio
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most TTS models take one sentence and read it. Sesame's CSM 1B takes the whole conversation so far — the actual audio of what was said, not just its transcript — and generates the next line grounded in that context, matching the pace, pitch, and emotional tone the conversation has already established. That distinction is the entire premise behind CSM, and it's what Sesame calls "voice presence" — the goal of making a spoken interaction feel real and understood rather than merely narrated.
What Is CSM 1B?
CSM (Conversational Speech Model) is an open-weight speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus CEO Brendan Iribe. Released as a 1-billion-parameter checkpoint in March 2025 under the fully permissive Apache 2.0 license, CSM generates speech directly from interleaved text and audio history — rather than from text alone — using a two-transformer architecture built on a LLaMA backbone.
Under the Hood: Two Transformers and the Mimi Codec
CSM's architecture splits the generation work across two LLaMA-style autoregressive decoders. A larger backbone decoder processes the interleaved text and audio context and predicts the first codebook token for the next piece of speech. A smaller depth decoder then takes that backbone's hidden state and generates the remaining acoustic codebook tokens needed to reconstruct fully intelligible audio.
Both decoders operate on tokens from Mimi, a split residual vector quantization (RVQ) audio codec developed by Kyutai. Mimi compresses speech down to roughly 1.1kbps at a 12.5Hz frame rate while preserving high audio fidelity — an efficient enough representation that CSM can run without the heavy computational demands more elaborate acoustic pipelines require.
The Real Differentiator: Grounding Speech in Actual Conversation
This is worth sitting with, because it's easy to undersell as a minor architectural detail. By giving the backbone decoder access to both past audio and past text simultaneously, CSM can ground its next utterance in the actual acoustic style of the conversation so far — including the speaker's pitch contour, pace, and emotional state — rather than generating a response purely from what the text says it should sound like. This is the specific mechanism behind CSM's ability to match a user's energy when they sound tired, sound more animated when they sound excited, and pause more naturally when a conversation's pace slows down. Most prior production voice systems generate from text representation alone, without that acoustic grounding in what actually just happened in the exchange.
An Honest Note: The Open Checkpoint vs. the Viral Demo
It's worth being precise about what you're actually getting. A fine-tuned variant of CSM powers Maya and Miles, Sesame's hosted voice companions whose demo went viral in February 2025 for sounding closer to a real conversational partner than prior commercial TTS systems. The publicly released CSM 1B checkpoint is the open-source base model underlying that experience — not the exact fine-tuned version behind the viral demo itself. Calibrate your expectations accordingly: the base model is a genuinely capable foundation, but the specific polish of the demo that made headlines involved additional fine-tuning Sesame hasn't released as open weights.
What CSM Can't Do
A few limitations are worth knowing before you build around it. CSM is trained specifically as an audio generation model, not a general-purpose multimodal LLM — it cannot generate text, so you'll need to pair it with a separate language model to produce the actual response content before CSM turns it into speech. It's also primarily an English-language model; it has some incidental capacity for other languages due to contamination in its training data, but that's not a capability to rely on for any serious non-English use case.
Sesame has also been explicit about acceptable use: the project's guidelines prohibit generating speech that impersonates a real individual without their explicit consent, creating deceptive or misleading content like fake news or fraudulent calls, and any illegal or harmful application of the technology.
Getting Started with CSM 1B
- Clone the repository.
git clone [email protected]:SesameAILabs/csm.git, then set up a Python 3.10 virtual environment and install dependencies viapip install -r requirements.txt. - Get access to both gated models on Hugging Face. You'll need approved access to both
sesame/CSM 1Bandmeta-llama/Llama-3.2-1B, since CSM's backbone depends on the Llama checkpoint — log in viahuggingface-cli loginbefore attempting to download either. - Disable lazy compilation in Mimi. Setting the environment variable
NO_TORCH_COMPILE=1is a documented step to avoid compilation issues specific to the Mimi codec component. - Use
triton-windowsinstead of standardtritonon Windows. The standardtritonpackage isn't installable on Windows, and the project's documentation specifically calls out this substitute. - Pair it with a separate LLM for full conversational pipelines. Since CSM only generates audio, plan for a text-generation model upstream in your pipeline to produce the actual dialogue content that CSM will then voice.
- Consider a community fork for a more complete out-of-the-box interface. Projects like
akashjss/sesame-csmadd a Gradio UI and an OpenAI-compatible API on top of the base model, with support across CUDA, MLX, and CPU devices — a faster path to a usable interface than building one from the base repository alone.
Tips for Better Results
- Feed it real conversation history, not just the line you want spoken. CSM's core advantage only shows up when you give it interleaved text and audio context to ground against — using it like a single-sentence TTS model without that history undersells what it's actually built for.
- Don't expect the exact quality of the viral Maya/Miles demo from the base checkpoint. That experience runs on a fine-tuned variant Sesame hasn't open-sourced; treat the public CSM 1B as a strong foundation to build and fine-tune from, not an exact replica of what went viral.
- Plan your pipeline around CSM being audio-only. Budget for a separate text-generation model feeding into CSM, rather than expecting it to handle dialogue generation and voicing in a single step.
- Stick to English for anything production-critical. Given the incidental, unreliable nature of its non-English capability, treat English as the only language this model is actually built and tested for.
- Respect the explicit usage guidelines around consent and impersonation. Given how convincingly context-aware CSM's output can be, take the project's own restrictions on non-consensual voice impersonation seriously rather than as boilerplate.
CSM 1B vs. Other Llama-Backbone Speech Models
| CSM 1B | Orpheus TTS | Typical reference-audio cloning models | |
|---|---|---|---|
| Core input | Interleaved text + audio conversation history | Text (plus optional voice reference) | Text + a single static reference clip |
| Backbone | Two-transformer LLaMA architecture | Llama-3B | Varies (diffusion, flow matching, etc.) |
| Codec | Mimi (Kyutai), 12.5Hz, 1.1kbps | SNAC | Varies |
| Can generate text | No, audio only | No, audio only | No |
| Context awareness | Core design goal | Not a primary focus | Not a primary focus |
| License | Apache 2.0 | Apache 2.0 | Varies |
CSM 1B's specific niche is genuine conversational grounding — matching the acoustic reality of an ongoing exchange rather than treating each line as an isolated generation task, which is a meaningfully different design goal than either narration-focused models or reference-clip voice cloning.
Frequently Asked Questions
What does CSM stand for?
Conversational Speech Model — reflecting its core design goal of maintaining contextual awareness across a dialogue rather than generating isolated, context-free utterances.
Can CSM 1B generate text responses on its own?
No. CSM is trained specifically as an audio generation model and cannot generate text — pair it with a separate language model to produce dialogue content before CSM converts it to speech.
Is the open-source CSM 1B the same model behind Sesame's viral voice demo?
Not exactly. A fine-tuned variant of CSM powers Maya and Miles, the hosted voice companions from Sesame's viral demo. The publicly released CSM 1B is the open-source base model underlying that system, not the specific fine-tuned version itself.
What is the Mimi codec?
Mimi is a split residual vector quantization audio codec developed by Kyutai, which CSM uses to encode speech into discrete tokens and decode them back into audio at roughly 1.1kbps while preserving high fidelity.
Is CSM 1B free for commercial use?
Yes, in terms of licensing — it's released under the Apache 2.0 license. That said, Sesame's usage guidelines explicitly prohibit non-consensual voice impersonation and deceptive use, which apply regardless of the permissive code license.
Turn a Multi-Speaker Script into Audio
If you already have a finished multi-speaker script, use Echora’s Text to Dialogue. Add up to 12 dialogue blocks with 5,000 characters in total, assign a voice to each block, add supported delivery tags, and adjust stability or speaker boost. You can then preview and download the complete conversation.