ElevenLabs Eleven v4 Launch
TECH

ElevenLabs Eleven v4 Launch

26+
Signals

Strategic Overview

  • 01.
    ElevenLabs launched Eleven v4, its most emotive text-to-speech model yet, and Eleven v4 Turbo, a low-latency variant for voice agents and real-time use, on September 28, 2026.
  • 02.
    Eleven v4 runs on an entirely new architecture that interprets tone, pacing, emotion, character, and context from text to generate expressive speech while keeping the speaker's identity consistent.
  • 03.
    Language support expanded from roughly 70 to more than 90 languages, with notable quality improvements in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.
  • 04.
    ElevenLabs markets Instant Voice Clones from just 10 seconds of audio, but its own v4 documentation says clones generally use one to two minutes of sample audio, a gap flagged as a best-case marketing claim.
  • 05.
    Eleven v4 Turbo achieves roughly 100ms median inference latency and about 150ms median time to first speech, outperforming Cartesia's Sonic 3.6 and OpenAI's GPT-4o mini TTS on that metric.
  • 06.
    Both models are available immediately in ElevenAgents, ElevenCreative, and via ElevenAPI, with access included on ElevenLabs' free account tier.
  • 07.
    Introductory launch pricing runs $22 per million characters for Eleven v4 and $11 per million characters for v4 Turbo over a two-week window.

A New Architecture, Not Just New Tags

Where Eleven v3's emotional control relied on explicit inline audio tags dropped into the text, Eleven v4 is built on an entirely new architecture that infers tone, pacing, emotion, character, and context directly from the words themselves, producing speech that can shift from dramatic to comedic to conversational while keeping a consistent speaker identity[1]. ElevenLabs says roughly 75% of listeners preferred v4 over competing models in blind head-to-head tests, though that figure comes from the company's own blog rather than an independent panel[1]. The company also touts Instant Voice Clones from as little as 10 seconds of audio, but its own v4 documentation describes clones as generally requiring one to two minutes of sample audio, a gap outside analysis flagged as a best-case marketing claim rather than the typical experience[2].

The Leaderboard Race: v4 vs Cartesia vs Google

Independent benchmarking firm Artificial Analysis put Eleven v4 at the top of its Provider Voice TTS Arena Leaderboard, with an Elo score of 1,319 across more than 1,600 head-to-head appearances, ahead of Cartesia's Sonic 3.6 (1,276) and Google's Gemini 3.8 Flash TTS (1,267)[3]. On the same benchmark suite, Eleven v4 also set a new high-water mark on Pronunciation Robustness at 91.7%, edging out Gemini 3.8 Flash TTS's 89.5%[3]. The picture is not uniformly dominant, though: on the separate Controlled Voice category, Eleven v4's Elo of 1,157 placed it second, behind Alibaba's Qwen-Audio-3.1-TTS-Plus. Runtime Wire cautions that a top leaderboard slot is evidence of a leading result in that specific comparison, not a universal measure of voice quality, since the arena tests native voices under blind pairwise conditions that may not match every real-world listening preference[2].

Funding the Race: $11B Valuation and Enterprise Revenue

The launch lands seven months after ElevenLabs raised $500 million led by Sequoia Capital in February 2026, more than tripling its valuation to $11 billion[4]. Co-founder Mati Staniszewski framed that round as fuel for the two product lines Eleven v4 now anchors, saying the company would expand its Creative offering by helping creators combine best-in-class audio with video and Agents[4]. The business context: ElevenLabs closed 2025 at $330 million in annualized recurring revenue and had grown that run rate past $600 million by the time of the v4 launch, with 55% of revenue now coming from large enterprise customers[5]. A model that reclaims the top leaderboard spot is also a data point for justifying that valuation to the market.

v4 Turbo Targets Live Voice Agents, With Real-World Stakes

Eleven v4 Turbo is the latency-optimized sibling, running at roughly 100ms median inference and about 150ms median time to first speech, beating Cartesia's Sonic 3.6 by 112ms and OpenAI's GPT-4o mini TTS by 664ms on time-to-first-speech[6]. ElevenLabs is explicit about where that speed is meant to go: ElevenAgents customer-support deployments handling confrontations, escalations, and holds, i.e. live phone conversations where a half-second of dead air breaks the illusion[5]. That positioning cuts both ways. Faster, more natural real-time voice agents mean AI can plausibly take over more live customer-service conversations that previously required a human on the line, and the same low-latency, highly expressive pipeline is just as usable for less benign real-time voice generation.

Not Everyone Is Sold: Cloning Caveats and Accent Homogenization

Reaction across early users was largely enthusiastic, with independent testers highlighting fine-grained control over both character performance and ambient soundscape, and multilingual clips that sound native across languages rather than obviously dubbed. But the skepticism clusters around specific, checkable claims. On voice cloning, the gap between ElevenLabs' promotional 10-second-of-audio figure and its own documented one-to-two-minute sample guidance is the clearest case of marketing outrunning the product spec[2]. Community testers also reported that some accented voices came out sounding more Americanized in v4 than in v3, a possible regression in accent fidelity even as headline benchmark scores improved, alongside broader doubts about whether voice consistency truly holds across the infinite text duration the company claims. None of that shows up in the leaderboard numbers, which measure blind pairwise preference, not longitudinal consistency or accent retention.

Historical Context

2023
Released Eleven Multilingual v2, extending language support to 28 languages, alongside lower-latency Flash v2.5 and Turbo v2.5 variants for live agent use.
2025-06-05
Released Eleven v3 in public alpha with 70+ languages, multi-speaker dialogue, and an audio-tag syntax for inline emotion/tone control.
2026-02-02
Eleven v3 reached general availability after roughly eight months in alpha.
2026-02-04
ElevenLabs raised a $500 million round led by Sequoia Capital, tripling its prior valuation to $11 billion.
2026-09-28
Launched Eleven v4 and Eleven v4 Turbo, its newest and most expressive text-to-speech models, topping the Artificial Analysis Provider Voice Arena leaderboard.

Power Map

Key Players
Subject

ElevenLabs Eleven v4 Launch

EL

ElevenLabs

AI voice/speech company that developed and launched Eleven v4 and v4 Turbo; competes on the Artificial Analysis Provider Voice TTS Arena leaderboard.

AR

Artificial Analysis

Independent benchmarking firm whose Provider Voice TTS Arena Leaderboard and Pronunciation Robustness benchmark ranked Eleven v4 #1, lending third-party credibility to ElevenLabs' quality claims.

CA

Cartesia (Sonic 3.6)

Competing voice-AI provider whose Sonic 3.6 model ranked #2 on the Provider Voice leaderboard and #1 on Controlled Voice, directly contesting ElevenLabs' claim of category leadership.

GO

Google (Gemini 3.8 Flash TTS)

Competing model that ranked #3 on Provider Voice and scored 89.5% on pronunciation robustness versus Eleven v4's 91.7%, representing large-platform competition in the TTS space.

SE

Sequoia Capital

Led ElevenLabs' $500M funding round in February 2026 that valued the company at $11 billion, giving it direct governance influence over ElevenLabs' product strategy including v4.

MA

Mati Staniszewski (ElevenLabs co-founder)

Company leadership driving product strategy that ties model releases like v4 to expansion of the Creative and Agents offerings.

Fact Check

6 cited
  1. [1] Eleven v4
  2. [2] ElevenLabs Eleven v4 Turbo Launch
  3. [3] ElevenLabs Eleven v4 Takes No. 1 Voice AI Rank With Pronunciation Record
  4. [4] ElevenLabs raises $500M from Sequoia at a $11 billion valuation
  5. [5] ElevenLabs' new v4 speech model supports more expression control and 90 languages
  6. [6] ElevenLabs Launches Eleven v4 With Low-Latency Turbo Variant

Source Articles

Top 4

THE SIGNAL.

Analysts

“Cautions that ElevenLabs' 'No. 1' leaderboard claim is not a universal quality measure and flags a discrepancy between the 10-second voice cloning marketing claim and official documentation.”

Runtime Wire (analysis outlet)
Industry commentary/analysis site

“Praised Eleven v4 Turbo for balancing conversational speed with quality, a gap prior real-time TTS models had not closed.”

Unnamed enterprise user (quoted by Unite.AI)
Customer/user of voice AI products
The Crowd

“Introducing Eleven v4 and Eleven v4 Turbo, our fastest and most emotive voice models yet. Ranked #1 by Artificial Analysis.”

@@ElevenLabs4516

“Truly blown away by the results from Eleven v4. This model has a new architecture that unlocks more realistic and controllable speech. You can now direct the performance of the character AND the soundscape around them. And the voice effects (like "cheap microphone") are on point.”

@@venturetwins371

“Same voice, four languages on Eleven v4 in this clip, and it sounds native in every one. So cool to see.”

@@n__deborah38

“ElevenLabs v4”

@u/aeroniero197
Broadcast
Introducing Eleven v4 and Eleven v4 Turbo

Introducing Eleven v4 and Eleven v4 Turbo

New Models and the Future of Voice AI | Mati Staniszewski, ElevenLabs Summit Warsaw 2026

New Models and the Future of Voice AI | Mati Staniszewski, ElevenLabs Summit Warsaw 2026