LAHORE – ElevenLabs has launched Eleven v4 and Eleven v4 Turbo, its latest text-to-speech models designed to deliver more expressive speech and faster responses.
The company announced the models in a post on X, describing them as its “fastest and most emotive voice models yet”.
The new models target different use cases. Eleven v4 focuses on speech quality and emotional delivery, while v4 Turbo is designed for applications that require faster responses.
Eleven v4 ranks highly in voice benchmark
ElevenLabs says Eleven v4 ranked first in the Artificial Analysis Provider Voice Arena in September 2026.
The company also conducted blind head-to-head preference tests against Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash-Lite TTS and Google Gemini 3.8 Flash TTS.
The tests focused on user preferences for expressiveness and naturalness. These results were reported by ElevenLabs and should be viewed in that context.
Eleven v4 focuses on expressive speech
Eleven v4 is positioned as the higher-quality model in the new lineup. The model is designed for content creation, audiobooks, character voices and dubbing, among other applications.
ElevenLabs says the model considers tone, pacing, emotion, character and context when generating speech.
It also supports inline audio tags such as [laughing], [whispering] and [shouting]. These tags give users more control over how generated speech is delivered.
Both Eleven v4 and v4 Turbo support more than 90 languages. They also offer multi-speaker dialogue and high-fidelity voice cloning. It additionally supports Professional Voice Clones and improved request stitching for longer content.
Eleven v4 Turbo targets real-time applications
Eleven v4 Turbo uses the same model family but focuses on lower latency. ElevenLabs says the model has a median inference latency of about 100 milliseconds, excluding application and network latency.
The company is positioning it for real-time applications such as AI assistants, customer support agents and interactive characters.
How the new model compares
ElevenLabs lists Eleven v4 Turbo at around 100ms median inference latency. By comparison, Eleven v3 Conversational has a median latency of about 280ms, while Eleven Flash v2.5 is listed at about 75ms.
Eleven v4 Turbo therefore sits between the two models in raw latency. However, ElevenLabs is positioning it as an option for users who want faster responses while retaining the expressive speech capabilities of the v4 family.
Availability
Eleven v4 and Eleven v4 Turbo are available through ElevenAgents, ElevenCreative and the ElevenAPI. The launch expands ElevenLabs’ AI voice lineup as developers increasingly use synthetic speech for assistants, customer service, entertainment and other interactive applications.
Key points:
- ElevenLabs has launched Eleven v4 and Eleven v4 Turbo, its latest text-to-speech models.
- The models focus on more expressive speech, improved voice cloning and faster response times.
- ElevenLabs says Eleven v4 ranked first in the Artificial Analysis Provider Voice Arena in September 2026.
- Eleven v4 Turbo records median inference latency of about 100 milliseconds.
- Both models support more than 90 languages, multi-speaker dialogue and high-fidelity voice cloning.
- Eleven v4 Turbo is designed for real-time applications such as AI assistants and customer support agents.