Skip to main content
Products & Tools

Eleven v4 and Eleven v4 Turbo Text to Speech models - ElevenLabs

Eleven v4 and Eleven v4 Turbo Text to Speech models &)]:tw-overflow-x-hidden"> strong]:tw-font-normal tw-text-gray-600"> Introducing Eleven v4 Meet Eleven v4, our most emotive model yet. With 3x credits included on Creator+ until October 12 Create controllable, expressive speech layered with emotion, audio events, and immersive soundscapes.

By Precis Daily Newsroom8 min read1,862 words
Illustration for: Eleven v4 and Eleven v4 Turbo Text to Speech models - Eleven
Illustration
Key points
  • Eleven v4 and Eleven v4 Turbo Text to Speech models &)]:tw-overflow-x-hidden"> strong]:tw-font-normal tw-text-gray-600"> Introducing Eleven v4 Meet Eleven v4, our most emotive model yet.
  • With 3x credits included on Creator+ until October 12 Create controllable, expressive speech layered with emotion, audio events, and immersive soundscapes.
  • Eleven v4 is built on an entirely new architecture that reads a script the way a voice actor would.

Eleven v4 and Eleven v4 Turbo Text to Speech models &)]:tw-overflow-x-hidden"> strong]:tw-font-normal tw-text-gray-600"> Introducing Eleven v4 Meet Eleven v4, our most emotive model yet. With 3x credits included on Creator+ until October 12 Create controllable, expressive speech layered with emotion, audio events, and immersive soundscapes. Eleven v4 is built on an entirely new architecture that reads a script the way a voice actor would. It knows who's speaking, what just happened, and how every line should land. John - host, theatrical [warm] Welcome back to the final round. [long pause] Our returning champion needs one more answer. [building] For ten thousand pounds: who wrote the Moonlight Sonata? [Gong sounds] Five seconds. Sia - conversational, British [whispered in a British accent] Beethoven. It's Beethoven, isn't it? [uncertain] I'm fairly sure it's Beethoven. [nervous laugh] Final answer. John - host, theatrical [delighted] Beethoven is correct! [crowd applause] [amused] I thought we'd lost you there for a moment. [warm] Ten thousand pounds! Speech in 90+ languages across an exceptional emotional range, with multiple speakers and sound effects built-in. Add direction like [laughs], [whispers], and [door slams] straight into the script. Eleven v4 follows tag sequences more reliably than v3, sound effects included. Context stitching keeps pacing and delivery steady across a script of any length. A full audiobook sounds like a single take from the first page to the last. Redo a line once or fifty times and it’s still the same person speaking. Speaker stability holds across dialogue, narration, and everything in between. Professional Voice Clones weren’t supported in v3. In Eleven v4, they’re back and perform with the model’s full emotional range across every language they speak. Introducing Eleven v4 Turbo for realtime and agents Our fastest real-time speech model with the expressive range of Eleven v4, available through the API and ElevenAgents. Eleven v4 Turbo has a median inference latency of ~100 ms, time to first speech of ~150 ms, and stays consistent in long interactions. At 150ms median time to first speech, the latency disappears and conversation flows smoothly. 150ms 1 0 1 2 3 4 5 6 7 8 9 1 5 0 1 2 3 4 5 6 7 8 9 5 0 0 1 2 3 4 5 6 7 8 9 0 ms 262ms 2 0 1 2 3 4 5 6 7 8 9 2 6 0 1 2 3 4 5 6 7 8 9 6 2 0 1 2 3 4 5 6 7 8 9 2 ms 814ms 8 0 1 2 3 4 5 6 7 8 9 8 1 0 1 2 3 4 5 6 7 8 9 1 4 0 1 2 3 4 5 6 7 8 9 4 ms Learn more about Eleven v4 Turbo from an agent powered by it. Push text as your LLM generates it and audio starts coming back before the sentence is finished. Bidirectional streaming, built for agent loops. Turbo carries the full expressive range of Eleven v4. Confirmations, escalations, and holds land differently from one another instead of reading identically. Professional Voice Clones work identically across both models, so a single brand voice stays consistent from the first turn of a call to the last. Point it at Japanese, Spanish or Portuguese text and the voice speaks it fluently, with a native accent. Narration voices A place for the storytellers. Warm, authoritative, and consistent across every project. Conversational voices For everyday conversations, you need a natural, easygoing voice. Our conversational voices are built for dialogue and podcasts that are made to feel unscripted. Social media voices Voices that come alive in short-form content. Capture the energy and personality needed to make a user stay on your TikTok or Reel from end to end. Character voices A pool of voices for your fictional world. Select from a dynamic cast that brings your game, audiobook, or animation to life. Educational voices Trustworthy, patient voices that guide students through complex topics. These voices excel for courses, tutorials, training simulations, and guides at every level. voices Polished voices brimming with confidence to land your next product sale. Perfect for digital ads or TV and radio reads. Entertainment voices From booming movie trailers to comic voices made for performance, these voices scream personality. Built for content where delivery matters as much as material. Multilingual voices Voices built for global content. Natural pacing, idiomatic expressions, and accents that connect with audiences in every market. We've scaled Agentforce Voice adoption through deterministic control, and what we hear consistently from customers is that they trust it because every action an agent takes is grounded and governed instead of improvised. A key part of this strategy is balancing that determinism with high-quality, high-EQ voice models. That's exactly where we're seeing Eleven v4 Turbo raise the bar, with faster, more natural responses that meet the standard our customers expect - allowing them to bring Agentforce Voice to even more use cases. Ryan Peterson , SVP Product for Agentforce Voice, Salesforce Working with hundreds of publishers to bring their journalism to audio, we see firsthand how much voice quality matters. Since partnering with ElevenLabs, many of our publishers have seen higher engagement and longer listening times. Eleven v4 gives publishers more ways to make sure those voices feel engaging, familiar, and distinctly their own. Patrick O'Flaherty , Co-Founder of BeyondWords ElevenLabs v4 brings a new level of natural sound to voice conversations. Voices are not only more expressive, but more controllable, giving us the ability to create richer, more immersive voice experiences. Chrys Bader , Co-founder and CEO of Rosebud First time I used it (Eleven v4), it was so clean and lifelike, it felt like I was running a session with talent in the booth. Kyle Gudmundson , Music & Audio Lead, Accenture Accelerate We've been looking for a voice model that's fast enough to feel like a real conversation without trading away quality, and Eleven v4 Turbo is the first one that does both. For automated sales workflows, this is the point where building stops feeling like an experiment. Oscar Daniels , Head of Credit Building Products, Spring Financial Choose how it sounds before it says a word Every generation starts with a voice. Clone one you already have, design one from a description, or correct how specific words are said with IPA support. Clone a voice from ten seconds of audio, or train a Professional Voice Clone for a near-perfect match. woman in her late 20s with a bright conversational voice and an American accent. Friendly, playful and relaxed tone. ” Describe a voice in a sentence and generate it, without a recording session or a casting call. Define how names, acronyms and technical terms are pronounced. Set phonetics ones, and every generation uses them. Start creating for free Upgrade when you need more :focus-visible]:tw-outline-focus tw-rounded-sm"> Get started for free Free plan comes with: Available via API Bring Eleven v4 to your app Build with Eleven v4 Turbo for real-time experiences and voice agents, all through ElevenAPI. Access the Eleven v4 and Eleven v4 Turbo API Add expressive speech to any product using our REST API, streaming endpoints, or TypeScript and Python SDKs. Stream text in and receive audio in real time Maintain continuity across long-form generations :focus-visible]:tw-outline-focus tw-isolate" style="box-shadow:0 0 1px rgb(0 0 0 / 0.4), 0 1px 1px rgb(0 0 0 / 0.04), 0 2px 4px rgb(0 0 0 / 0.04);border-top-left-radius:1.25rem;border-top-right-radius:1.25rem;border-bottom-right-radius:1.25rem;border-bottom-left-radius:1.25rem"> import { ElevenLabsClient, play } from '@elevenlabs/elevenlabs-js' ; const elevenlabs = new ElevenLabsClient({ const audio = await elevenlabs.textToSpeech.convert( 's3TPKV1kjDlVtZbl4Ksh' , // "George" - browse voices at elevenlabs.io/app/voice-library text: 'The first move is what sets everything in motion.' , Yes. Instant Voice Clones in Eleven v4 now outperform the Professional Voice Clones of Multilingual v2. Professional Voice Clones are back after being unavailable in v3, and perform with the model's full emotional range. Every clone requires verified consent from the voice's owner. Every voice in the voice library (17,500+ voices) works with Eleven v4. For PVCs and IVCs created prior to Eleven v4's launch, to work effectively you will need to retrain them with Eleven v4. This can be done by hovering over the voice and clicking the plus button. How many languages does Eleven v4 support? Eleven v4 supports 90+ languages, allowing you to create rich audio samples for audiences across the world. A single generation supports up to 10,000 characters. For long-form content, context stitching keeps pacing and delivery consistent across generations. What's new in Eleven v4 compared to Eleven v3? Eleven v4 is built on a new architecture with higher audio quality and a wider emotional range. Speaker identity is now stable across regenerations. Professional Voice Clones are supported again, Audio Tag following is more reliable, and time to first byte has dropped. What's the difference between Eleven v4 and Eleven v4 Turbo? Eleven v4 is tuned for produced content where quality matters most. Eleven v4 Turbo is a low-latency variant at ~100 ms median inference latency, designed for voice agents and real-time use. Both models share the same expressive range and both support Professional Voice Clones. What audio quality and formats does Eleven v4 output? Eleven v4 outputs to the same formats as every ElevenLabs model, so it drops into any existing workflow: MP3: Standard format for podcasts, YouTube, and general listening. WAV / PCM: Uncompressed audio for studio work, dubbing, and post-production. µ-law: Optimized for telephony and call-center integrations. Sample rate and bitrate are set via the API, so Eleven v4 audio can be tuned for quality or bandwidth depending on where it's going. How does ElevenLabs handle data privacy and security with Eleven v4? Eleven v4 runs on the same infrastructure and compliance posture as every ElevenLabs model, trusted by leading enterprise customers: Scripts and audio tags you send to Eleven v4 are not used to train our models without your consent. Enterprise customers can enable Zero Retention Mode for eligible services.* Every voice clone in Eleven v4, instant or professional, requires verified consent from the voice's owner, and all generated audio is protected by AI Speech Classifier technology that can detect it as AI-generated. * For ZRM-eligible services, where ZRM is correctly enabled, certain types of data are not retained. See documentation for details. How do I control pauses, emphasis, and pronunciation in Eleven v4? Inline audio tags are the control mechanism in Eleven v4. Write [pause] or [long pause] where you want a break, and tags like [whispers], [excited], or [sighs] to shape delivery. SSML tags such as are disabled in Eleven v4, so use the natural-language tags instead. Pronunciation dictionaries still define how names and technical terms are spoken. Yes. Both streaming and non-streaming endpoints launch alongside the model, with TypeScript and Python SDKs. Eleven v4 Turbo is designed for the agent loop, with median ~100 ms inference latency and bidirectional streaming. It's available through ElevenAgents and the API. How much does Eleven v4 cost? Is there a free plan? Eleven v4 uses the same credit pricing as our other Text to Speech models, so it's available on every plan including the free tier, which includes 10,000 credits a month, or roughly 10 minutes of audio.

Sources

Summarized from the linked originals.

Related stories

Illustration for: The latest AI news we announced in September 2026
Models & Research

The latest AI news we announced in September 2026 Here’s a recap of some of our biggest AI updates from September, including Gemini 4 Argon, new Connected Apps in the Gemini App, and WeatherNext 3. Your browser does not support the audio element.

Google AI Blog8 min