ElevenLabs ships v4 voice models, ranked No. 1 by Artificial Analysis

The new models add performance controls and a real-time version; a two-week introductory API offer starts at $22 per million characters.

By · Published

Primary source: ElevenLabs on X

Why it matters

ElevenLabs is selling the same speech research into creator software and live customer-facing agents. The leaderboard supports its quality claim in one defined test, while the gap between its 10-second cloning claim and its documentation's usual sample guidance shows why customers still need to validate the model against their own workflows.

A high-end studio microphone sits on a desk, with abstract, colorful sound wave visualizations flowing on a blurred background monitor.

ElevenLabs co-founders Mati Staniszewski (@mati) and Piotr Dabkowski (@dabkowski_piotr) have released Eleven v4 and Eleven v4 Turbo, a pair of speech-generation models built to make AI voices easier to direct and faster to use in live conversations. The company announced the models in a September 28th post on X, saying they are available through ElevenAgents, ElevenCreative and ElevenAPI.

https://x.com/ElevenLabs/status/2104572127617994917

poster=/api/storage/public-objects/tweet-videos/elevenlabs-eleven-v4-turbo-launch-poster-f665cf4f.jpg|Video from @ElevenLabs on X

The launch splits one product bet across two kinds of work. Eleven v4 targets creators making dialogue, narration and character performances, with inline directions for emotion, pacing, reactions and sound effects. The Turbo variant is aimed at real-time conversations. ElevenLabs says Turbo has a median inference latency of about 100 milliseconds, a company-reported figure that matters most in applications where pauses can make a voice agent feel unresponsive.

The release also puts a measurable claim behind ElevenLabs' quality pitch. Artificial Analysis' Provider Voice Arena leaderboard lists Eleven v4 at No. 1, with an Elo score of 1,319 from 1,674 samples. That is evidence of a leading result in that specific comparison, not a universal measure of voice quality: the leaderboard compares models using each provider's native voices, and listeners may value qualities beyond the arena's ranking. ElevenLabs' post cites the ranking without spelling out the methodology.

Staniszewski and Dabkowski started ElevenLabs in 2022. Dabkowski previously worked on machine learning at Google and now leads research and engineering at the company; Staniszewski is CEO. Their original focus on generated speech has grown into products for creators, developers and enterprise voice agents. A February Series D announcement valued ElevenLabs at $11 billion after a $500 million round led by Sequoia Capital, with Andreessen Horowitz and ICONIQ among the significant follow-on investors. ElevenLabs said in May that it had exceeded $500 million in annual recurring revenue during the first four months of 2026; that is a company-reported run rate, not audited revenue.

That commercial footprint gives the model launch a practical test: can better-controlled speech make automated conversations sound natural enough for customer-facing work, while staying responsive? ElevenLabs is positioning v4 Turbo for support, sales and scheduling agents, while the standard v4 model reaches creator workflows. The same underlying voice technology can therefore serve both applications, and the company is packaging it across its own tools and developer API.

There is a detail buyers should check before building a workflow around the launch copy. The X thread says Instant Voice Clones can capture a voice from 10 seconds of audio. ElevenLabs' v4 documentation says an Instant Voice Clone generally uses a sample of one to two minutes. The discrepancy makes the 10-second claim worth treating as a best-case claim until users test it with their own recordings. The documentation also says v4 handles cross-language cloning differently: when generated speech is in a different language from the source recording, the model aims for a fluent target-language accent rather than retaining the source speaker's accent by default.

ElevenLabs says v4 supports more than 90 languages, including Cantonese, Mongolian and Odia, and adds International Phonetic Alphabet support for pronunciation control. The API launch offer runs for two weeks: Eleven v4 at $22 per million characters and Turbo at $11 per million characters. ElevenLabs also says v4 is included at no additional cost for eligible Creator+ plans, subject to monthly credit limits.

The promotional pricing and broad availability make this a product launch as much as a model release. ElevenLabs is lowering the barrier for developers to try the new voices while steering customers toward its own creative and agent products. Whether the improved expressiveness holds up across languages, voices and extended conversations will determine how much the leaderboard position translates into production use.

Reader comments

Conversation for this story loads after sign-in.