ElevenLabs Now Copies a Voice From 10 Seconds of Audio, in More Than 90 Languages

ElevenLabs released Eleven v4 on Monday. It’s a speech model the company says is built for audiobooks, character work, voiceovers and dubbing, and it can make an instant copy of a voice from 10 seconds of audio, then have that voice speak more than 90 languages. Last week Google asked for 30 seconds and a spoken yes. Today’s sample is a third as long, and the launch post doesn’t mention consent at all.

Two models, one very short sample

ElevenLabs put out two models on September 28: Eleven v4, which it calls its most emotive yet, and Eleven v4 Turbo, a faster version for voice agents that talk to customers in real time. Both are live now in the company’s creator app, its agent platform and its developer API.

The line that matters most for working voice actors sits near the bottom of the announcement. “Instant Voice Clones can now capture voices with high fidelity using just 10 seconds of audio,” the company writes. The model also “adds support for Professional Voice Clones (PVC), for the highest-fidelity cloning use cases.”

Ten seconds is shorter than most commercial reads. It’s shorter than the opening spot on a typical demo reel. Anyone with a public demo, a YouTube channel or a podcast guest slot has already published that much audio many times over.

TechCrunch’s Ivan Mehta reports that the language count rises from 70 in the previous model to more than 90, and that ElevenLabs saw its biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese. Those are four of the busiest dubbing markets in the world.

It’s pitched as a performer, not a reader

Read the announcement as a voice actor and it sounds like a casting brief. The models are “built for content where delivery matters as much as the words themselves, whether that’s audiobooks, character performances, voiceovers, dubbing, or localizing conversational agents,” ElevenLabs says.

The example it chooses is pure acting class. A line like “I need you to stay calm” should sound different depending on who says it, the post explains, “whether that’s a doctor delivering it gently to a frightened patient, or a character in a game shouting to his squad before dropping into battle.”

Direction now arrives inside the script. Users type tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing], and the company says “Eleven v4 follows these audio tags and direction prompts more accurately than prior models, so the delivery you describe is what you get.” That’s the job of a director and an actor in a booth, written into a text box.

Infographic: what Eleven v4 changed for voices, covering the 10-second sample, more than 90 languages, cross-language clones, direction tags and consent

One voice, every language

The feature with the sharpest edge for dubbing is the cross-language clone. “Now, a voice recorded in one language also speaks any other fluently, adopting the accent of a native speaker while retaining the identity of the original,” ElevenLabs says. In Martin Cid Magazine, Susan Hill puts it plainly: “a Spanish speaker’s clone reading German still sounds like that person, with a native German accent.”

VoiceEditSuitePro audio tools built for working voice actors.Auto Clean UpACX Audiobook PrepAudition Director$20/mo$15/monthShow me moreCancel anytime. No contract.

For a studio, that means one English session could in theory become the Spanish, Portuguese and Japanese tracks too. For the actors who used to voice those tracks, it’s the scenario they’ve warned about all year. Fabio Azevedo, president of the Brazilian Association of Dubbing Professionals and the Brazilian voice of Doctor Strange, told Rest of World in April: “We make foreign content sound Brazilian with our Brazilian idiosyncrasies; with AI, we lose that.” Brazilian Portuguese is one of the four languages ElevenLabs singles out as most improved.

Ten seconds, and the consent question

The launch post says nothing about consent. The company’s voice cloning page does: “You can freely clone your own voice for any purpose. Cloning someone else’s voice requires their explicit consent.” Unite.AI’s launch report notes that professional clones require verified consent from the voice’s owner.

Google, September 23

Gemini 3.8 Flash TTS

A 30-second sample. Before a clone is made, the voice’s owner records a spoken consent statement, and the system checks it matches the sample.

ElevenLabs, September 28

Eleven v4

A 10-second sample for an instant clone. The cloning page says someone else’s voice needs their explicit consent. The launch post doesn’t raise it.

That’s a different shape of safeguard from the one Google announced last week. As Voice Over Herald reported, Google’s new speech models ask the owner to say yes out loud on a recording the system compares with the sample. A matching voice is harder to fake than a tick box.

Writing for developers on the OrcaRouter blog, Gideon Frost argues the shorter sample raises the stakes rather than lowering them. “The consent problem does not shrink with the sample. It grows, because the bar to produce a convincing clone has fallen to a clip anyone can record,” he writes. If you clone a voice you don’t own, he adds, “the compliance question is now entirely yours to answer upstream of the API call,” because “the model will do what you ask.”

Hill flags the risk outside the studio too: “a scammer needs only a short sample of a relative’s voice to stage a fake emergency call.”

The money behind the model

ElevenLabs is not a small lab experimenting on the side. TechCrunch reports that its annualized revenue run rate has climbed from roughly $330 million at the start of the year to over $600 million, that it raised $500 million from Sequoia at an $11 billion valuation, and that “more than 55% of its business” now comes from large companies. Chief executive Mati Staniszewski told the outlet the company is aiming for an IPO “in the next years.”

Listeners are a different story. In July the Audio Publishers Association’s executive director, Jim Dinegar, told Publishers Weekly that “the human voice is strongly preferred by audiobook listeners.” The same report put AI-narrated titles at 0.03% of audiobook sales revenue in 2025.

What working voice actors can do now

Read the AI clause before you sign. Ganessh Divekar, general secretary of the Association of Voice Artists of India, described the new habit to Rest of World: “Now, we are telling members to ask what it will be used for, and ask for more money when signing perpetuity contracts.” With a 10-second threshold, a perpetuity clause on a single tag line is worth reading twice.

Know where your audio lives. Every public demo, reel and interview is now long enough to clone. That’s not a reason to hide your work. It’s a reason to know where it is, so you can spot a copy when one turns up.

Watch Tokyo tomorrow. On September 30 a Tokyo court is due to rule in Kenjiro Tsuda’s case against TikTok over an AI copy of his voice, the first ruling of its kind in Japan. Whatever the judges decide about one famous voice will shape how far a performer can go when a stranger copies theirs.

The technology is now built to cast, direct and localize in one pass. The part it can’t supply is the part consent protects: a real person who agreed to be heard that way.

Sources: TechCrunch (Ivan Mehta, September 28, 2026); ElevenLabs, “Introducing Eleven v4, our most emotive model” and the ElevenLabs voice cloning page; Unite.AI; Martin Cid Magazine (Susan Hill); OrcaRouter blog (Gideon Frost); Rest of World (Rina Chandran, April 15, 2026); Publishers Weekly (Ed Nawotka, July 14, 2026); Voice Over Herald. Hero: composite image, The Voice Realm.

Lee este artículo en español: ElevenLabs ahora copia una voz con 10 segundos de audio, en más de 90 idiomas

Know something about this story, or spotted an error? Tell the newsroom.