A New Architecture, Not Just New Tags
Where Eleven v3's emotional control relied on explicit inline audio tags dropped into the text, Eleven v4 is built on an entirely new architecture that infers tone, pacing, emotion, character, and context directly from the words themselves, producing speech that can shift from dramatic to comedic to conversational while keeping a consistent speaker identity[1]. ElevenLabs says roughly 75% of listeners preferred v4 over competing models in blind head-to-head tests, though that figure comes from the company's own blog rather than an independent panel[1]. The company also touts Instant Voice Clones from as little as 10 seconds of audio, but its own v4 documentation describes clones as generally requiring one to two minutes of sample audio, a gap outside analysis flagged as a best-case marketing claim rather than the typical experience[2].

