YFarmX logoYFarmX

ElevenLabs

ElevenLabs

the voice infrastructure standard

Released 2 February 20264 min readAudio and VoiceLast updated:

Editorial illustration: ElevenLabs

Key facts

90+Eleven v4
Languages
75msFlash v2.5
Latency
150msScribe v2 Realtime
Transcription
#1Eleven v4, 1320 Elo
TTS arena
LicensedMerlin, Kobalt
Music data
v2.5default, 11 Sep 2026
Music model

The default API for synthetic voice, tuned for very low latency and built on licensed music data.

Eleven v4 is the current generation of the speech model that has made ElevenLabs the default supplier of synthetic voice on the web. The company released it on 28 September 2026 alongside a low-latency variant, Eleven v4 Turbo, and both cover more than 90 languages, up from the 70-plus of Eleven v3, which they replace, according to the launch post. Eleven v4 follows expressive audio tags, markers such as [laughs] or [phone buzzing] that let a script call for a whisper, a laugh or a sound effect rather than leaving delivery to chance, and ElevenLabs says it follows them more accurately than earlier models. Turbo answers in a median of about 150 milliseconds to first speech, and Instant Voice Clones now need 10 seconds of audio, both on the company’s own figures. On Artificial Analysis’ text-to-speech arena, read on 2 October 2026, Eleven v4 is first on 1320 Elo, ahead of Qwen-Audio-3.1-TTS-Plus on 1291.

Built for low latency

Voice is only useful in an application if it arrives quickly, and much of the ElevenLabs range is tuned for latency. Flash v2.5 reaches roughly 75ms latency, fast enough for conversational agents where any audible lag breaks the illusion of a real exchange. On the input side, Scribe v2 Realtime targets around 150ms for speech to text, closing the loop so an agent can listen and respond in something close to natural time. Together these define the platform as infrastructure for live voice rather than a tool for pre-rendered clips. The numbers are the point: a delay a human would barely notice in a recording becomes glaring in a two-way conversation, so shaving latency to the tens of milliseconds is what turns a demo into a usable agent.

The move into music

The company has also moved into generated music. ElevenMusic arrived with prompt-to-music with vocals, inpainting to regenerate a chosen section, and multilingual output, which places it alongside the dedicated music models even though voice remains the core business. The current model is Music v2.5, published on 11 September 2026 and now the default for prompted and reference-based generation in ElevenMusic, per the company’s announcement. In a blind test across 47,885 paired comparisons, listeners preferred the v2.5 take the majority of the time, with the widest gap on vocal-led and acoustic-heavy genres such as R&B, soul, hip hop, rock, metal, orchestral and cinematic music, on the company’s own figures. The line’s distinguishing feature is provenance: ElevenMusic was trained exclusively on licensed material from partners including Merlin Network and Kobalt Music Group. That gives it the cleanest commercial position for background tracks, an advantage that becomes concrete the moment a customer needs music they can use in a product or a campaign without inheriting a rights dispute. For a producer, that provenance is the difference between a track that is safe to publish and one that carries hidden risk.

Why provenance wins buyers

Provenance is the thread running through the whole proposition. Where several rivals in generative audio carry unresolved questions about training data, a licensed foundation lets ElevenLabs sell to businesses that cannot take on legal exposure. For an advertiser, a game studio or a corporate producer, being able to point to named licensing partners is often the deciding factor, and it complements the technical lead that Eleven v4 represents on the voice side. It is a commercial argument as much as a technical one, and for many buyers it settles the choice before a single sample is played.

Why teams pick it

The result is a platform that has been the top-ranked AI voice service through mid-2026 and the default API for teams building voice agents, audiobooks and accessibility narration. Each of those use cases leans on a different strength: agents need the low latency of Flash and Scribe, audiobooks need the expressive range of Eleven v4, and accessibility narration needs both reliability and broad language coverage, which the 90-plus languages supply. Bundling all of it behind one API is what has made ElevenLabs the standard rather than one option among many.

Where it sits and what to watch

The placement is straightforward: ElevenLabs occupies the infrastructure layer of AI audio, the part other products build on rather than compete with directly, and it anchors our AI audio coverage. What to watch next is whether the music push gains ground against specialists, and whether the low-latency, licensed model holds its lead as larger platforms fold real-time voice into their own stacks. Eleven v4 is the benchmark others are measured against, first on the independent arena at its launch, and the licensed foundation beneath ElevenMusic is as important as the voices themselves. For the broader context, see our AI hub.