ElevenLabs has released Eleven v4, a speech synthesis model that follows direction cues more accurately and keeps voices stable across long productions. The launch also includes v4 Turbo for real-time voice agents, which starts producing speech in about 150 milliseconds. For business users, this reduces the trade-off between expressive voice quality and response speed in customer-facing automation.

ElevenLabs releases Eleven v4 with faster Turbo for voice agents

Direction tags, consistency and Turbo latency

Eleven v4 generates laughter, whispers and non-speech sounds such as slamming doors more reliably than its predecessor. Eleven v3, released just over a year ago, already supported such audio tags but followed them less accurately. Users can direct delivery through tags or plain sentences, and can use phonetic spelling to set pronunciation of names and technical terms. The company says pronunciation controls now work more reliably.

Narrators and characters are expected to sound consistent throughout a production, even when users regenerate individual lines several times. A single request handles up to 10,000 characters, roughly ten minutes of audio, while longer works like audiobooks are assembled from multiple segments with consistent pacing across transitions. In dialogue, AI speakers respond to the context of the entire scene rather than delivering each line in isolation.

The model supports more than 90 languages, up from about 70 in v3. Cloned voices are expected to keep native accents in other languages without drifting back to the original accent over time. Professional Voice Clones are supported again after being unavailable in v3, while Instant Voice Clone requires only ten seconds of audio. ElevenLabs licenses voices from the people behind them, including celebrity voices such as Michael Caine through a dedicated marketplace.

What this means for voice automation buyers

For companies running service calls, sales calls and in-product assistants, Turbo targets the combination of expression and low latency. ElevenLabs reports 150 milliseconds to audible speech for Turbo, against 262 milliseconds for Cartesia Sonic 3.6 and 814 milliseconds for OpenAI GPT-4o mini TTS. Turbo was optimized together with the ElevenAgents platform, and both models are available in ElevenAgents, ElevenCreative and through the API. Small teams gain faster deployment of natural-sounding agents, while larger contact centers can test expressive voices without a separate latency penalty.

Quality claims still require independent verification before procurement decisions. Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google Gemini 3.8 Flash TTS on the Artificial Analysis Provider Voice Arena leaderboard, and scores 91.7 percent on pronunciation against 85.6 percent for v3. In company blind tests, about three-quarters of listeners preferred v4 over models from Cartesia, Inworld and Google. Buyers should test accent stability, pronunciation of industry terms and consistency after line regeneration on their own scripts.

The marker to watch is pricing after October 12. Standard API pricing is $80 per million characters for v4 and $40 for Turbo, falling to $22 and $11 through that date, while Artificial Analysis lists Sonic 3.6 at $49 and Gemini 3.8 Flash TTS at $16.49. Users on the $22 monthly Creator plan or higher can use v4 in ElevenCreative at no extra cost for two weeks, capped at twice monthly credits. Whether list prices hold will show how ElevenLabs values quality against cheaper high-volume alternatives.