ElevenLabs
ElevenLabs
the voice infrastructure standard

Key facts
- 90+Eleven v4
- Languages
- 75msFlash v2.5
- Latency
- 150msScribe v2 Realtime
- Transcription
- #1Eleven v4, 1320 Elo
- TTS arena
- LicensedMerlin, Kobalt
- Music data
- v2.5default, 11 Sep 2026
- Music model
The default API for synthetic voice, tuned for very low latency and built on licensed music data.
Eleven v4 is the current generation of the speech model that has made ElevenLabs the default supplier of synthetic voice on the web. The company released it on 28 September 2026 alongside a low-latency variant, Eleven v4 Turbo, and both cover more than 90 languages, up from the 70-plus of Eleven v3, which they replace, according to the launch post. Eleven v4 follows expressive audio tags, markers such as [laughs] or [phone buzzing] that let a script call for a whisper, a laugh or a sound effect rather than leaving delivery to chance, and ElevenLabs says it follows them more accurately than earlier models. Turbo answers in a median of about 150 milliseconds to first speech, and Instant Voice Clones now need 10 seconds of audio, both on the company’s own figures. On Artificial Analysis’ text-to-speech arena, read on 2 October 2026, Eleven v4 is first on 1320 Elo, ahead of Qwen-Audio-3.1-TTS-Plus on 1291.
Built for low latency
Voice is only useful in an application if it arrives quickly, and much of the ElevenLabs range is tuned for latency. Flash v2.5 reaches roughly 75ms latency, fast enough for conversational agents where any audible lag breaks the illusion of a real exchange. On the input side, Scribe v2 Realtime targets around 150ms for speech to text, closing the loop so an agent can listen and respond in something close to natural time. Together these define the platform as infrastructure for live voice rather than a tool for pre-rendered clips. The numbers are the point: a delay a human would barely notice in a recording becomes glaring in a two-way conversation, so shaving latency to the tens of milliseconds is what turns a demo into a usable agent.
The move into music
The company has also moved into generated music. ElevenMusic arrived with prompt-to-music with vocals, inpainting to regenerate a chosen section, and multilingual output, which places it alongside the dedicated music models even though voice remains the core business. The current model is Music v2.5, published on 11 September 2026 and now the default for prompted and reference-based generation in ElevenMusic, per the company’s announcement. In a blind test across 47,885 paired comparisons, listeners preferred the v2.5 take the majority of the time, with the widest gap on vocal-led and acoustic-heavy genres such as R&B, soul, hip hop, rock, metal, orchestral and cinematic music, on the company’s own figures. The line’s distinguishing feature is provenance: ElevenMusic was trained exclusively on licensed material from partners including Merlin Network and Kobalt Music Group. That gives it the cleanest commercial position for background tracks, an advantage that becomes concrete the moment a customer needs music they can use in a product or a campaign without inheriting a rights dispute. For a producer, that provenance is the difference between a track that is safe to publish and one that carries hidden risk.
Why provenance wins buyers
Provenance is the thread running through the whole proposition. Where several rivals in generative audio carry unresolved questions about training data, a licensed foundation lets ElevenLabs sell to businesses that cannot take on legal exposure. For an advertiser, a game studio or a corporate producer, being able to point to named licensing partners is often the deciding factor, and it complements the technical lead that Eleven v4 represents on the voice side. It is a commercial argument as much as a technical one, and for many buyers it settles the choice before a single sample is played.
Why teams pick it
The result is a platform that has been the top-ranked AI voice service through mid-2026 and the default API for teams building voice agents, audiobooks and accessibility narration. Each of those use cases leans on a different strength: agents need the low latency of Flash and Scribe, audiobooks need the expressive range of Eleven v4, and accessibility narration needs both reliability and broad language coverage, which the 90-plus languages supply. Bundling all of it behind one API is what has made ElevenLabs the standard rather than one option among many.
Where it sits and what to watch
The placement is straightforward: ElevenLabs occupies the infrastructure layer of AI audio, the part other products build on rather than compete with directly, and it anchors our AI audio coverage. What to watch next is whether the music push gains ground against specialists, and whether the low-latency, licensed model holds its lead as larger platforms fold real-time voice into their own stacks. Eleven v4 is the benchmark others are measured against, first on the independent arena at its launch, and the licensed foundation beneath ElevenMusic is as important as the voices themselves. For the broader context, see our AI hub.
More in Audio and Voice
All Audio →- Alibaba QwenQwen-Audio-3.1Alibaba's upgraded audio stack, announced 23 September 2026
- NvidiaNemotron 3 DiarizationNvidia’s Nemotron 3 Diarization marks when each speaker is active in recorded or live audio
- Microsoft AIMAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flashlive transcription that tops the streaming accuracy board, and two new voices
- AlibabaHappyShrimp 1.0the name is a eulogy
- Microsoft AIMAI-Transcribe-2speech recognition at ten cents an hour
- SunoSuno v6a flagship, an exploratory variant and a free mini, released together