ElevenLabs released Eleven v4 and its real-time sibling Eleven v4 Turbo on September 28, 2026, with 90+ languages, more reliable inline audio tags, and the return of Professional Voice Clones. Turbo starts speaking in about 150ms — more than five times faster than OpenAI's GPT-4o mini TTS in ElevenLabs' September measurements — and the family topped the Artificial Analysis Provider Voice Arena.
On September 28, 2026, ElevenLabs launched Eleven v4 and Eleven v4 Turbo, the next generation of its text-to-speech lineup. The flagship model is built on an entirely new architecture that, in the company's words, 'reads a script the way a voice actor would' — knowing who is speaking, what just happened, and how every line should land. Both models support speech in more than 90 languages with native accents, and they are available through the ElevenCreative web app, the API with REST and streaming, and the ElevenAgents conversational platform, with a launch promotion of 3x credits on Creator+ plans until October 12.
Eleven v4 is aimed at audio you produce once and publish, such as audiobooks and video voice-overs. Inline audio tags like [laughs], [whispers], and [door slams] are written directly into the script, and v4 follows tag sequences more reliably than v3, sound effects included. Context stitching keeps pacing and delivery steady across a script of any length so a full audiobook 'sounds like a single take,' while speaker stability means regenerating a line dozens of times still produces the same person speaking. Professional Voice Clones — absent from v3 — are back and perform with the model's full emotional range, instant clones still work from a 10-second sample, and the model works with more than 17,500 existing voices, outputting MP3, WAV, or PCM plus mu-law for telephony at up to 10,000 characters per generation.
Eleven v4 Turbo is built for real-time voice agents and telephony. ElevenLabs reports a median inference latency of about 100ms and a median time to first speech of about 150ms, against 262ms for Cartesia Sonic 3.6 and 814ms for OpenAI's GPT-4o mini TTS in the company's own September 2026 measurements over WebSockets with identical scripts — a gap of more than five times against OpenAI. Turbo supports bidirectional streaming, so audio can start before the LLM finishes a sentence, and carries the full expressive range of v4. ElevenLabs says v4 ranked first on the third-party Artificial Analysis Provider Voice Arena leaderboard for September 2026, and that roughly 75% of listeners preferred it in the company's own blind tests against five rival services. Salesforce's Ryan Peterson, SVP Product for Agentforce Voice, said v4 Turbo 'raises the bar' with 'faster, more natural responses' for grounded, governed voice agents.
The two-model split — expressive studio quality for produced audio, low latency for live conversation — mirrors how the market is separating audiobook- and media-grade TTS from agent-grade real-time speech. BeyondWords co-founder Patrick O'Flaherty said publishers using ElevenLabs have seen 'higher engagement and longer listening times,' and that v4 gives them more ways to keep voices 'engaging, familiar, and distinctly their own,' while Rosebud CEO Chrys Bader called the models 'not only more expressive, but more controllable.' With voice clones back, 17,500+ voices, and a five-fold latency lead over OpenAI's GPT-4o mini TTS in the company's measurements, Eleven v4 positions ElevenLabs across both the studio and the call center.