Resemble AI: Voice Cloning Built With a Security Team's Instincts
- What Is Resemble AI?
- Functional Advantages of Resemble AI
- How Resemble AI Voice Cloning Works
- How Resemble AI's Security Stack Actually Works
- Getting Started with the Resemble AI API
- Tips for Better Resemble AI Results
- Resemble AI vs. ElevenLabs
- Frequently Asked Questions
- Create a Similar Voice from Reference Audio
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
Most voice cloning tools treat security as an afterthought — a terms-of-service line item, not a product feature. Resemble AI built its business the other way around: it clones voices from a short sample like any other platform, but every clip it generates carries an invisible watermark, and the same company sells the detection technology used to catch deepfakes made by other tools. If your use case involves brand voices, talent likenesses, or any content where "prove this is authentic" matters as much as "make it sound real," Resemble AI's combination of cloning and provenance is worth understanding in detail.
What Is Resemble AI?
Resemble AI is a voice AI company founded in 2019 by Zohaib Ahmed and Saqib Muhammad, now headquartered in the San Francisco Bay Area after starting in Toronto. The company has raised roughly $25 million across five funding rounds, including a $13 million strategic round in December 2025 backed by Google AI Futures Fund, Sony Innovation Fund, KDDI, and Okta Ventures — investors whose own interest in fraud prevention and telecom security says as much about Resemble's positioning as its product pages do. Its customers span entertainment studios, Fortune 500 telecom providers, and government agencies, and the company operates as two connected businesses under one platform: a generative voice studio, and a deepfake-detection and watermarking suite it also sells standalone. Resemble AI is also the creator of Chatterbox, the MIT-licensed open-source TTS model many teams use as a self-hosted alternative to commercial cloning tools.
Functional Advantages of Resemble AI
- Clone a voice from a short sample, then keep going from there — Resemble builds a usable voice clone from roughly 30–60 seconds of source audio, with a higher-fidelity "Pro" clone tier for production use.
- Every generated clip is watermarked by default — Resemble's PerTH watermarking embeds a perceptually hidden, tamper-resistant signal into each output, so provenance can be verified later even after re-encoding or format conversion.
- A deepfake-detection engine most competitors don't have at all — Resemble Detect (built on models like DETECT-3B Omni and DETECT-World) analyzes audio, video, and image content frame-by-frame in real time to flag synthetic or manipulated media, sold both as an add-on to voice generation and as a standalone security product.
- Pay-per-use pricing with no forced monthly tier — the Flex plan bills by the second across TTS, voice cloning, and detection, with credits that never expire, instead of locking you into a subscription sized for someone else's usage pattern.
- Consent verification built into the cloning flow — cloning someone else's voice requires confirmation of consent at creation time, backed by the same watermarking and detection stack used to catch unauthorized use after the fact.
How Resemble AI Voice Cloning Works
Upload a clean reference recording — as little as 30 to 60 seconds for a usable clone — and Resemble generates a voice model you can drive with any text input, adjusting pacing and emotional delivery as you go. A lower-cost "Rapid" clone tier suits quick prototypes and one-off projects, while the "Pro" clone tier costs more per month but produces the higher-fidelity result needed for a recurring brand voice or a licensed talent likeness. Real-time speech-to-speech conversion is also available, letting a live speaker's voice be converted to the target voice as they talk rather than generating from text.
How Resemble AI's Security Stack Actually Works
This is the part of the platform that sets Resemble apart from most voice-cloning competitors, and it runs on two separate technologies that are designed to work together:
Watermarking (PerTH) embeds an inaudible signal directly into the audio waveform at generation time. The watermark is built to survive common transformations — compression, re-encoding, even being used as training data for another model — so a later scan can confirm a clip originated from Resemble even after it's been re-uploaded, edited, or stripped of metadata.
Detection (Resemble Detect) is the inverse capability: instead of marking content Resemble generated, it analyzes content from any source to flag whether it was likely AI-generated or manipulated. The current detection models — DETECT-3B Omni and the newer DETECT-World — run frame-by-frame analysis across audio, video, and images, checking both known synthesis patterns and physical-consistency signals like lighting and continuity, which is what lets DETECT-World catch output from generators it wasn't specifically trained on. Detection runs as a parallel process over an API, so it doesn't add latency to a live call or stream it's monitoring.
Together, watermarking and detection form what Resemble calls a "trust stack" — provenance you can prove for your own content, and a way to check content you didn't generate. For teams cloning talent voices, brand voices, or anything with legal exposure attached, this is the practical answer to "how do we prove this audio is (or isn't) ours."
Getting Started with the Resemble AI API
- Create an account and open the API page. Sign up at app.resemble.ai, then go to Account → API to find your unique API token.
- Copy your API key. Resemble displays your token directly on that page rather than a one-time reveal — store it as an environment variable rather than hardcoding it into a project.
- Install the SDK for your language. Official libraries are available via
pip install resemble(Python),npm install @resemble/node(Node.js), andbundle add resemble(Ruby). - Choose a request type. Resemble's API supports three synthesis modes — Async, Sync, and Streaming — so you can pick immediate playback, batch generation, or continuous streaming depending on the use case.
- Make your first call. A typical flow authenticates with
Resemble.api_key('YOUR_API_KEY'), retrieves a project and voice UUID, then calls the clips endpoint with your text to generate audio.
Tips for Better Resemble AI Results
- Record your source sample in a quiet space with a consistent mic distance. Even at 30–60 seconds, background noise is the most common cause of a clone that sounds inconsistent across longer scripts.
- Use the Pro clone tier for anything recurring. The Rapid tier is fine for a one-off test, but a brand voice or ongoing character deserves the higher-fidelity option.
- Model your usage before committing to Flex at scale. Per-second billing is cheap to start but can outpace a flat subscription on high-volume workloads like a busy IVR — estimate your monthly seconds before assuming it's the cheaper option.
- Turn on Detect if you're publishing to public channels. Pairing watermarking on your own output with detection on incoming or third-party content closes the loop rather than only covering half of it.
- Get documented consent before cloning anyone else's voice. Resemble's own terms of service require it, and the watermarking and detection stack exists partly to enforce that after the fact.
Resemble AI vs. ElevenLabs
| Resemble AI | ElevenLabs | |
|---|---|---|
| Pricing model | Pay-per-second, no monthly minimum | Flat monthly subscription tiers |
| Voice cloning | ~30–60 sec sample, Rapid/Pro tiers | Instant (1–3 min) / Professional (30 min–3 hrs) tiers |
| Built-in watermarking | Yes — PerTH, applied by default | Limited, not a core product focus |
| Deepfake detection | Yes — Resemble Detect, sold standalone | Not offered as a product |
| Sound effects / dubbing | Not a core offering | Yes — dedicated sound effects and dubbing tools |
| Best fit | Security- and provenance-sensitive use cases | Broad content production across TTS, cloning, and audio |
Neither platform is strictly better — they're optimized for different priorities. ElevenLabs bundles more creative tools (sound effects, dubbing, a large voice library) into one flat-rate subscription. Resemble AI narrows its focus to cloning plus the trust layer around it, which matters most if your voice content needs to be provably authentic — brand voices, licensed talent, or anything where a synthetic audio dispute could become a legal question.
Frequently Asked Questions
What is Resemble AI used for?
Resemble AI is used for voice cloning, real-time speech-to-speech conversion, brand and character voice creation, and deepfake detection or content authentication through its Detect and watermarking tools.
Is Resemble AI free?
The Flex plan starts at $0 with pay-as-you-go pricing and no monthly commitment — you're only billed for what you generate, per second, plus a small monthly fee for each voice clone you keep active.
How does Resemble AI's watermarking work?
Resemble embeds a perceptually hidden, tamper-resistant watermark (PerTH) into every clip it generates. The watermark is designed to survive re-encoding and format changes, so provenance can be verified later even if the file has been edited or re-uploaded elsewhere.
How much source audio does Resemble AI need to clone a voice?
As little as 30 to 60 seconds of clean reference audio is enough to produce a usable clone, with a higher-fidelity Pro tier available for production or brand-voice use.
Does Resemble AI require consent to clone a voice?
Yes. Cloning someone else's voice without their permission violates Resemble AI's terms of service, and the platform requires confirmation of consent during the clone-creation process.
How do I get a Resemble AI API key?
Log in at app.resemble.ai, go to Account → API, and copy the token shown there — it works across Resemble's Async, Sync, and Streaming endpoints.
Create a Similar Voice from Reference Audio
For a browser-based workflow with a similar goal, use Echora’s Voice Clone. Upload a recording of your own voice or one you have explicit permission to use, enter up to 5,000 characters of new text, and generate new speech that resembles the reference voice.