ElevenLabs v4, 10-Second Voice Cloning

ElevenLabs released its v4 and v4 Turbo voice models on September 29, supporting more than 90 languages, 10-second audio cloning, and finer-grained emotional control. The Turbo version has a median time to first speech of approximately 150 milliseconds, clearly targeting real-time voice agents.
ElevenLabs v4: Clone a Voice in 10 Seconds
ElevenLabs released two new voice models on September 29: ElevenLabs v4 and Eleven v4 Turbo. The former targets high-quality content production such as audiobooks, dubbing, and character performances, while the latter focuses on voice agents and real-time calls.
The core of this update is not simply that “the voice sounds more human.” ElevenLabs is attempting to address three long-standing problems in speech generation at the same time: voice identity tends to drift in long-form text, emotional control is not granular enough, and models take too long to start speaking during real-time conversations. The v4 series responds with support for more than 90 languages, instant voice cloning from just 10 seconds of audio, and a new architecture plus expressive tags for controlling tone, pacing, and emotion.
For developers, the Turbo version is the one that deserves the most attention. According to data released by ElevenLabs, v4 Turbo has a median time to first audio of approximately 150 milliseconds, significantly below the 262-to-814-millisecond range commonly seen in some high-quality speech models. This difference may not seem substantial for a single line of narration, but in continuous conversations involving customer service, sales, medical consultations, or game NPCs, the experience can shift from “voice playback” to something much closer to real-time communication.

From “Being Able to Speak” to “How to Speak”
Traditional text-to-speech models are generally good at reading text aloud, but they do not necessarily understand what emotion should be used. The sentence “Okay, I’ll help you take care of that” could express confirmation, perfunctory agreement, reassurance, or a customer-service agent attempting to calm a customer before escalating a complaint. For humans, context determines tone. For models, this has long been a dividing line in generation quality.
ElevenLabs introduced inline tags in v3, allowing users to specify emotions and speaking styles directly in the text. In v4, tag-based control has been enhanced further: developers can stack multiple tags, and the model executes them in the order they appear. For example, a line of dialogue can first request a quiet delivery and a pause, then switch to a more resolute tone. This kind of control is closer to the way a director gives instructions to a voice actor than a simple “emotional intensity” slider.
More importantly, v4 does not process only individual sentences in isolation. The model considers the surrounding textual context, adjusts its tone across longer passages, and attempts to preserve the speaker’s vocal identity. For audiobooks, educational content, and continuous narratives, this means developers no longer need to repeatedly slice text, tune parameters, and stitch large numbers of audio clips back together.
This also signals a shift for voice models from “sentence-level generation” toward “paragraph-level performance.” A genuinely usable dubbing system does not merely generate every sentence clearly enough. It must make listeners feel that the same person is speaking continuously, with emotions changing according to the context while the speaker’s identity remains consistent.
Ten-Second Cloning Lowers the Trial Barrier
ElevenLabs says that v4’s Instant Voice Clone can create a high-fidelity voice clone from just 10 seconds of audio. In the past, obtaining relatively stable cloning results generally required longer, cleaner recordings, and sometimes even a professional voice-cloning process. The significance of a 10-second sample is that voice capture can be embedded into more product workflows: a user records a short clip and can then obtain personalized narration, virtual characters, or multilingual content.
However, “cloneable in 10 seconds” does not mean that “10 seconds is enough to train a professional voice actor.” Short samples are better suited to quickly capturing vocal identity. The final result is also affected by the recording environment, spoken content, vocal-range coverage, and pronunciation consistency. For formal commercial projects, companies will still need higher-quality source recordings, as well as review and revision of the generated output.
Another improvement in v4 is the preservation of vocal identity across languages. Users can record their voice in one language and have the model generate speech in other supported languages, using a native accent for the target language where possible while avoiding a gradual drift back toward the original language’s accent later in the output. ElevenLabs specifically highlighted noticeable improvements in synthesized Japanese, Brazilian Portuguese, Mandarin, and Cantonese.
The value for content localization is direct. In the past, when a brand voice entered different language markets, companies often had to find new voice actors or accept a compromise in which the voice sounded similar but was clearly not the same person. v4 aims to turn brand voices, character voices, and personal voices into assets that can be reused across languages.
Of course, lowering the barrier to voice cloning also increases the risk of abuse. ElevenLabs says that both instant and professional voice cloning in v4 require consent verification from the voice owner. Generated audio is also protected by its AI Speech Classifier technology to help identify AI-generated content. For developers, this means voice authorization, identity verification, generation records, and complaint handling cannot be treated as patches added after launch. They should be incorporated into the product design from the beginning.
Turbo Targets the Last Mile of Voice Agents
The experience of a real-time voice agent is usually determined by several stages working together: the user’s speech must be recognized, a large language model on the backend must generate a response, and the voice model must convert that response into audio. If any one of these stages slows down, the user will hear a noticeable silence.
v4 Turbo is designed to compress this waiting time as much as possible. ElevenLabs says that when the backend large language model has just begun producing a response, Turbo can generate speech simultaneously instead of waiting for the entire answer to be completed before synthesizing it all at once. This streaming process is similar to downloading and playing a video at the same time: the system speaks the content that is already available, then continues with subsequent content as it arrives.
In customer-service and phone scenarios, this mechanism is more valuable than simply improving audio quality. Customers may not require every sentence to sound like a movie performance, but they are highly sensitive to long silences, mechanical repetition, and irrelevant responses. Turbo’s differentiated handling of scenarios such as disputes, escalation requests, and call waiting suggests that ElevenLabs is putting its voice model into real business workflows rather than treating it merely as a text-to-audio API.
It may be suitable for the following scenarios:
- Customer service and after-sales support: Maintain the conversation with a natural waiting tone while retrieving an order or calling a business system, instead of leaving users in silence.
- Sales assistants: Adjust speaking speed, pauses, and vocal emphasis according to the customer’s emotions, reducing the “same tone for every sentence” problem common in traditional voice bots.
- Medical consultations: Conduct preliminary Q&A with a gentler, steadier tone while maintaining accurate pronunciation of specialized terminology.
- Games and virtual characters: Allow characters to change their speaking speed, emotions, and speaking styles as the story develops instead of merely playing prerecorded audio.
- Multilingual services: Use the same brand or character voice across different regions while adapting to the accent and communication habits of the target language.
How to Choose Between v4 and v4 Turbo
The two models are not simply higher- and lower-spec versions. They are better understood as two versions designed for different workflows.
| Model | Primary positioning | Best suited for | Main trade-off | | --- | --- | --- | --- | | ElevenLabs v4 | High-quality generation with strong expressiveness | Audiobooks, dubbing, long-form text, character performances | Prioritizes final audio quality and emotional fidelity | | ElevenLabs v4 Turbo | Low-latency real-time generation | Voice agents, phone calls, game interactions, real-time applications | Prioritizes response speed and continuous conversation quality |
If the task is to generate several minutes of narration in one pass, or to finely tune emotion, character, and paragraph-level pacing, v4 is the better fit. If users may interrupt, ask follow-up questions, or change topics at any point during generation, Turbo’s response latency becomes more important.
This also means developers do not necessarily need to choose a single model for the entire product pipeline. Turbo can handle real-time conversations, while v4 can be used for post-call summaries, training recordings, marketing content, or high-quality replays. As long as voice identity, output format, and invocation methods remain consistent, the product can switch models based on the scenario without redesigning the entire voice workflow.
Its Competitive Edge Is Not Just Audio Quality
ElevenLabs is no longer competing only with traditional TTS providers. Cartesia, PlayHT, Microsoft, Google, and numerous startups building around voice agents are all competing for the same market: making AI not merely capable of “answering,” but able to carry out a conversation like a real service representative.
ElevenLabs’ advantage lies in its established voice assets and creator-product foundation. Its voice library contains more than 17,500 voices and covers multiple scenarios, from content creation to enterprise calls. By placing expressiveness, multilingual support, and low latency into the same update, the new models show that the company is shifting its focus from “does the voice sound good?” to “can the voice become a deployable business component?”
However, engineering trade-offs still exist between low latency and high expressiveness. Real-time systems must consider more than time to first audio. They also need to account for long-running stability, interruption handling, streaming-audio jitter, cost, and peak concurrency. A model that starts speaking within 150 milliseconds in a demonstration does not necessarily maintain the same speed throughout a complex business workflow.
When evaluating v4 Turbo, developers should test at least four things:
- After a user interrupts, can the model stop the current audio promptly and move into the next turn?
- During streaming output of a long response, do voice identity, volume, and emotion remain stable?
- Do mixed Chinese and English, numbers, brand names, and industry terminology produce abnormal pronunciations?
- Under high concurrency, do latency, failure rates, and unit audio costs remain acceptable?
ElevenLabs Is Moving Closer to Becoming an Infrastructure Company
This release also comes as ElevenLabs’ commercial expansion accelerates. The company says that more than 55% of its revenue now comes from large enterprises, its annualized revenue run rate has grown from approximately $330 million at the beginning of the year to more than $600 million, and its headcount has surpassed 800 employees. It is also continuing to hire in markets including India, Europe, and Brazil.
Earlier this year, ElevenLabs raised $500 million from Sequoia Capital at a post-money valuation of $11 billion. Recent market reports have also suggested that the company is preparing for another funding round, potentially targeting a valuation of $22 billion. Co-founder and CEO Mati Staniszewski has also said that the company plans to complete an IPO within the next few years, although no specific timeline has been announced.
The underlying change represented by these figures is that voice generation is shifting from a content tool into enterprise infrastructure. Companies are no longer purchasing merely the ability to “read text aloud.” They are purchasing multilingual voice assets, real-time conversation, telephone-system integration, compliance auditing, and stable API services.
For developers in China, real-world deployment also requires consideration of network connectivity, regional availability, data compliance, audio formats, and cost control. Teams that need to integrate GPT, Claude, Gemini, DeepSeek, and other models alongside voice capabilities can use an OpenAI-compatible aggregation platform such as OpenAI Hub to manage model calls centrally, reducing the cost of repeatedly adapting to different vendor SDKs. However, streaming interfaces, audio formats, and real-time communication capabilities for voice models still need to be validated separately. Text-model compatibility claims alone are not enough.
Assessment: v4’s Value Lies in Making Voice Programmable
The most important change in ElevenLabs v4 is not that it supports more than 20 additional languages, nor that it reduces voice cloning to 10 seconds. It is that voice is becoming more programmable: developers can specify the order of vocal styles, the model can retain context, cloned voices can be reused across languages, and Turbo can bring these capabilities into real-time interaction.
This will change how voice applications are developed. In the past, voice was often an output layer added at the end of the product process. Going forward, voice may become part of the product logic, much like visual themes, character settings, and dialogue strategies.
However, this is still far from saying that “voice agents are solved.” Natural-sounding speech is only one layer of the experience. Whether the model is accurate, can call business systems, and avoids making reckless promises in high-risk scenarios will still determine whether a product can truly go live. v4 and v4 Turbo address expression and response time; developers still need to solve permissions, fact verification, human escalation, and auditing.
As of September 29, ElevenLabs had integrated the new models into ElevenAgents, ElevenCreative, and ElevenAPI, and free accounts could also access them. For teams building multilingual content, real-time customer service, or character-based applications, this update is worth testing directly with real business audio. If the task is only one-off dubbing, v4’s improvements may be noticeable enough, but they may not justify immediately rebuilding an existing workflow.
Sources
-
IT Home: ElevenLabs Launches the v4 and v4 Turbo Voice Models — Comprehensive reporting on the model release date, language support, 10-second voice cloning, enterprise business data, and funding information.
-
Official ElevenLabs v4 Introduction (IT Home report) — Product information on emotion control, real-time voice agents, and multilingual generation in v4 and v4 Turbo.



