LOVO AI (Genny): Voiceover, Video, and Voice Cloning in One Browser Tab
Turn text and recordings into audio you can hear
Dubbing, stems, covers, and sound effects — all in the browser.
- Text to speech, plus voice cloning
- Split stems, or generate a cover
- Sign up to preview — listen first, decide later
LOVO AI is a Berkeley, California-based company founded in 2019 by Charlie Choi and Tom Lee, which raised a $4.5 million Pre-Series A round in August 2021 led by Kakao Entertainment and LG CNS. Its product, Genny, set out to solve a specific workflow problem: producing a finished, voiced-over video usually means bouncing between a text-to-speech tool, a video editor, and a subtitling tool, exporting and re-importing files at every step. Genny bundles all of it — script writing, voice generation, video editing, and auto-subtitles — into a single browser-based workspace, positioned for social media creators, marketers, and e-learning teams who'd rather not manage three separate subscriptions for one finished video.
Functional Advantages of LOVO AI (Genny)
- Voice generation and video editing share one workspace. A generated voiceover drops directly onto a timeline alongside footage and music, instead of being exported from one tool and re-imported into another.
- A large, broad voice library. 500+ voices across 100+ languages and 30+ preset emotional styles cover a wide range of accents and use cases without needing to clone anything.
- Pro V2's natural-language direction is a genuinely different interface. Instead of nudging sliders for pitch, speed, and emphasis, you write bracketed instructions like
[sobbing],[british accent], or[speak more slowly and sound nervous]directly into the script, and the model interprets the direction. - Voice cloning from a short sample. A custom voice can be built from roughly one minute of recorded audio, reused across unlimited projects on plans that include the feature.
- Auto-generated subtitles tied to the same voice generation call. Captions are produced from the same script and timing data as the voiceover, rather than requiring a separate transcription pass.
- An AI script writer and art generator in the same tab. For short-form content, a rough script and background visuals can be produced without leaving the workspace, even before voice generation starts.
How Genny's Directable Voices and Editor Actually Work
The more interesting technical piece here is Pro V2, launched in May 2025, which reframes voice direction as a language problem rather than a parameter-tuning problem. Older TTS interfaces — including LOVO's own earlier voices — expose emotion and delivery as separate sliders or presets you adjust outside the script. Pro V2 instead accepts instructions written inline, in square brackets, immediately before the text they apply to: [laughing and joking around] hahaha, and he said what? or [super embarrassed] I can't believe I just did that. The model reads the bracketed direction the way an actor would read a stage direction, adjusting delivery, pacing, and non-verbal sounds like laughter or sobbing to match — and it can generate multiple takes of the same line, letting you pick the best read rather than accepting the first output.
The video side of Genny is built around the same "single call" principle: because the voiceover, its timing data, and the script all originate from one generation step, auto-subtitles can be derived directly from that data instead of running a separate speech-to-text pass afterward, and the timeline editor can place the generated audio precisely against footage without a manual sync step. That's the practical value of the all-in-one architecture — it's less about any single feature being best-in-class and more about removing the export-and-realign friction that comes from stitching together a TTS tool, a video editor, and a subtitle generator built by three different companies.
Tips for Better Results with Genny's Voices
- Write bracketed directions as you would a stage direction, not a mood label.
[speak warmly, as if reassuring a nervous client]gives the model more to work with than a single word like[nervous], since it's interpreting natural language rather than selecting from a fixed emotion list. - Generate multiple takes on lines that carry real emotional weight. Pro V2 supports multiple takes per line specifically because delivery can vary between generations — treat the first result as a draft, not a final answer, on lines where tone matters most.
- Test voice quality across several of the 500+ voices before committing to a project. Reviewers consistently note more quality variation across LOVO's library than on more tightly curated platforms, so the "best" voice for your content isn't guaranteed to be an early or default pick.
- Record cloning source audio in a quiet room with a decent mic, even though one minute is technically enough. As with any short-sample cloning method, background noise in that one minute gets baked directly into the resulting voice.
- Build in buffer time around rendering, especially for longer 1080p exports. Cloud rendering speed varies with server load, and longer video exports have been reported to take noticeably longer during peak usage periods.
LOVO AI (Genny) vs. Descript Overdub
| LOVO AI (Genny) | Descript Overdub | |
|---|---|---|
| Core workflow | Script-to-voice-to-video in one workspace | Transcript-based editing of existing recordings |
| Voice cloning | ~1 minute, any voice you have consent for | Only your own voice, consent-gated by design |
| Video handling | Built-in timeline editor for new video creation | Edits existing video/audio via its transcript |
| Best fit | Building a video from a script and voiceover from scratch | Fixing or narrating over content you've already recorded |
| Emotional direction | Bracketed natural-language cues (Pro V2) | Boundary-matched insertion into real takes |
Both tools bundle voice generation into a broader editing environment rather than shipping it as a standalone feature, but they're solving different production stages. Genny is built for creating a video from nothing — script, voice, visuals, and captions generated together. Overdub is built for correcting or narrating over audio you've already recorded, and restricts cloning to your own voice specifically to avoid the consent questions that are central to LOVO's current legal situation. If your workflow starts from a blank page, Genny's all-in-one approach is the more direct fit; if it starts from an existing recording that just needs a fix, Overdub solves a narrower, more specific problem.
Frequently Asked Questions
Is LOVO AI (Genny) still operating?
As of this writing, Lovo Inc. has filed for Chapter 7 bankruptcy (a liquidation filing) while facing a class-action lawsuit over alleged unauthorized voice cloning, and the case has been stayed as a result. The website remains accessible, but service reliability and account access have been inconsistently reported — verify current status before relying on it for paid or production work.
What is LOVO AI (Genny) used for?
Genny is used for producing voiced-over videos end to end — script writing, text-to-speech narration, voice cloning, video editing, and auto-subtitle generation — inside one browser-based workspace.
How does LOVO AI's voice cloning work?
A custom voice can be created from roughly one minute of recorded audio, after which the resulting voice model can be reused across generations on plans that include the feature.
What are Pro V2 voices?
Pro V2 is LOVO's directable text-to-speech engine, launched in May 2025, which accepts natural-language instructions written in square brackets — like [sobbing] or [british accent] — directly inside a script, rather than requiring separate emotion or pacing controls.
How many voices and languages does LOVO AI (Genny) support?
The library spans 500+ voices across 100+ languages, with 30+ preset emotional styles on standard voices in addition to Pro V2's natural-language direction.
Create a Voiceover for Your Video with Echora
For a similar script-to-voice workflow in your browser, use Echora's Text to Speech. Enter up to 5,000 characters, choose a voice, and generate your narration. Preview the result and download an MP3 to use in your video editor.