Games · Model Watch
Eleven v4 puts more performance into AI voices
ElevenLabs is pushing text-to-speech beyond simply reading lines correctly. V4 is designed to understand how those lines should actually be performed.
Reporting updated Sep 29, 2026
From the creator
ElevenLabs’ official launch post for Eleven v4 and Eleven v4 Turbo.View original post on X ↗
THE TAKEAWAY
ElevenLabs is pushing text-to-speech beyond simply reading lines correctly. V4 is designed to understand how those lines should actually be performed.
What Happened
ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28, with the new generation focused heavily on expressive performance. V4 is designed to interpret context such as tone, pacing, emotion and character rather than treating each line as an isolated piece of speech. ElevenLabs says the model can handle delivery ranging from dramatic or urgent to comedic and conversational while maintaining the speaker’s identity. Multi-speaker conversations have also been improved so dialogue can react more naturally across speakers instead of sounding like separately generated lines stitched together. The model adds support for Professional Voice Clones and more reliable stitching for longer-form generations. The release supports more than 90 languages. ElevenLabs is also promoting v4 Turbo for lower-latency uses such as interactive characters and voice agents, reporting roughly 100 milliseconds of median inference latency. Those latency figures are company measurements rather than an independent VEYR test. Both models are available through ElevenCreative, ElevenAgents and ElevenAPI.
Why It Matters
Character voice is one of the places where an otherwise impressive AI production can still fall apart fast. A technically clean voice is not necessarily a convincing performance. Actors react, pause, interrupt, change intensity and interpret what a line means in context. That makes v4 particularly interesting for game dialogue, animated characters, temp performances, localization and fast iteration on scenes where delivery matters. Turbo potentially extends the same idea into interactive characters, where the response cannot arrive several seconds after the player speaks. The open question is consistency. A polished company demo can show the model at its best. Creators will need to find out how reliably that performance control survives longer scenes, repeated characters and dozens or hundreds of lines.
VEYR Angle
Voice generation is starting to look less like text-to-speech and more like performance direction. That changes the creator’s job too. Instead of only writing what the character says, creators increasingly need to describe how the character feels, where the energy shifts and how one speaker reacts to another. For film and game creators, that is a much more useful evolution than simply making synthetic speech sound cleaner.
.webp)