YFarmX logoYFarmX

Microsoft AI

MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash

live transcription that tops the streaming accuracy board, and two new voices

Released 1 October 20266 min readAudio and Voice

Editorial collage headed Microsoft MAI, with a black headset on a desk whose microphone spills a curling paper transcript and whose earcup sends out blue sound waves, the Microsoft logo, and tags reading #1 accuracy, 23 languages and $0.54 an hour; the subtitle reads live transcription and Voice 2.1.

Key facts

1 Oct 2026Microsoft AI, public preview
Released
2.51% WERfirst of 38, Artificial Analysis
Streaming accuracy
$0.54 / hourintroductory, to 31 Dec 2026
Transcription price
$22 / $15per M characters, 2.1 / 2.1-Flash
Voice prices
23one voice keeps its identity
Voice languages
60with automatic detection
Transcription languages

Microsoft AI released three speech models on 1 October 2026 for building voice agents that listen and talk in real time. MAI-Transcribe-2-Streaming turns live audio into text as people speak, in 60 languages, and ranks first for accuracy on Artificial Analysis's streaming board at a 2.51% word error rate. MAI-Voice-2.1 and MAI-Voice-2.1-Flash turn text into speech in 23 languages, at $22 and $15 per million characters, with Flash built for fast replies in live conversation.

Microsoft AI released three speech models on 1 October 2026, all aimed at voice agents: software that listens to a person, decides what to do and answers out loud, fast enough to feel like a conversation. MAI-Transcribe-2-Streaming turns live speech into text as it is spoken. MAI-Voice-2.1 and MAI-Voice-2.1-Flash turn text into speech, in 23 languages.

“A voice agent is a loop. It has to hear, understand, decide, and speak,” the launch post says. “And do it all within the window where a human still experiences the interaction as a conversation.” All three run through Azure Speech in Microsoft Foundry, in public preview.

MAI-Transcribe-2-Streaming writes while you talk

MAI-Transcribe-2-Streaming transcribes audio as a continuous stream, returning partial transcripts that update while the speaker talks and final transcripts that confirm each segment, in 60 languages with automatic language detection. It is the streaming sibling of MAI-Transcribe-2, the batch model Microsoft AI released on 3 September 2026, and its first partials arrive “in just over 100ms of receiving audio”, Microsoft says.

Developers connect through the Azure Speech SDK or over a WebSocket protocol compatible with OpenAI’s Realtime API. On the WebSocket route they send 16-bit mono audio in small chunks, and each session can last up to an hour, according to Microsoft Learn. Word-level timestamps, keyword biasing and speaker labels stay with the batch model, Microsoft’s comparison table shows.

It ranks first for accuracy on the streaming board

Artificial Analysis ranks MAI-Transcribe-2-Streaming first of 38 streaming models for accuracy, at a 2.51% word error rate on final transcripts and 2.52% on first partials, on its board read on 2 October 2026. The index mixes three test sets, around eight hours of voice-agent speech, European Parliament proceedings and company earnings calls, according to its methodology. Artificial Analysis reported it “taking the #1 spot” on 1 October.

Model, rank on accuracy Word error rate, final Time to final transcript Price per 1,000 minutes
MAI-Transcribe-2-Streaming, 1st 2.51% 0.128 s $9.00
Grok Voice Transcribe 2.0, 2nd 2.73% 0.490 s $3.33
Muse Voice Transcribe, 3rd 3.06% 0.163 s $3.00
Cartesia Ink Preview (external endpoints), 4th 3.11% 0.105 s $4.00
ElevenLabs Scribe v2 Realtime, 7th 3.59% 0.141 s $6.50
GPT Live Transcribe, 9th 3.92% 0.812 s $17.00
Artificial Analysis scatter chart of streaming speech-to-text word error rate, from 0 to 12%, against seconds to the final transcript, from 0 to 1.4, with a shaded most attractive quadrant, a dotted Pareto line and about 32 labelled models coloured by vendor. MAI-Transcribe-2-Streaming is the brown dot at the low-error end of the Pareto line, at 2.51% and 0.13 seconds.
Word error rate against time to the final transcript for 32 of the 38 models on Artificial Analysis's streaming board, with MAI-Transcribe-2-Streaming at the most accurate end. Select the chart to enlarge. Chart: Artificial Analysis, 2 October 2026.

Accuracy is the claim it wins. On the same board ten models return a final transcript sooner, led by Deepgram Flux at 0.021 seconds, and its $9.00 per 1,000 minutes is one of the higher prices: 28 of the 38 models cost less. Microsoft says it sits “on the Pareto frontier” for accuracy against latency.

Transcription costs $0.54 an hour

MAI-Transcribe-2-Streaming costs $0.54 per hour of audio, an introductory price that runs to 31 December 2026, according to its model card. That is 5.4 times the $0.10 an hour, also an introductory price, that Microsoft charges for the batch MAI-Transcribe-2, which waits for a recording to end before it transcribes.

Model What it does Price
MAI-Transcribe-2-Streaming Live speech to text $0.54 per hour of audio
MAI-Transcribe-2 Recorded speech to text $0.10 per hour of audio
MAI-Voice-2.1 Text to speech, most expressive $22 per M characters
MAI-Voice-2.1-Flash Text to speech, low latency $15 per M characters

At those rates, an hour of a voice agent listening costs $0.54 to transcribe, and 100,000 characters of spoken reply cost $1.50 on Flash.

Two voices that keep their identity across 23 languages

MAI-Voice-2.1 and MAI-Voice-2.1-Flash speak 23 languages, and one voice keeps the same identity in each of them with a native accent, Microsoft says. “Just ask it to speak English… then Mandarin… then German… and the speaker stays unmistakably the same,” the launch post says. Microsoft Learn lists 97 prebuilt voices, 50 male and 47 female, all usable with both models, and styles such as joy, excitement and empathy set through speech markup.

MAI-Voice-2.1 MAI-Voice-2.1-Flash
Built for Audiobooks, voice-overs, long-form speech Call centres, voice assistants, live dialogue
Model inference latency, Microsoft About 550 ms About 45 ms
Price per million characters $22 $15

Flash streams its output, so a caller can interrupt it mid-sentence. The quality figures are Microsoft’s own. In blind side-by-side comparisons over 5,032 judgments, listeners preferred MAI-Voice-2.1 with new prompts 58.9% of the time against MAI-Voice-2 with old ones. In a test with 4,000 listeners, 50.3% rated the MAI voices as human-like as or more human-like than human recordings, with 2.1 and 2.1-Flash results combined, the model page says. Flash delivers “55% faster model inference” and is about 60% cheaper than comparable models, Microsoft adds, in a comparison against ElevenLabs’ published model documentation.

The MAI Playground Chatter screen: a sphere drawn in dots above the heading Chatter, the line Talk to a voice assistant powered by our transcription and voice models, and a Start a conversation button, with a sidebar listing MAI-Transcribe-2, MAI-Voice-2.1, MAI-Image-2.6 and MAI-Thinking-1 and a Chatter Beta entry.
Chatter, the MAI Playground demo that runs on the new transcription and voice models. Source: Microsoft AI, 2 October 2026.

Voice cloning is gated. Microsoft’s documentation describes “instant voice cloning” from a 5 to 60 second reference clip, available only after approval through Azure’s limited access review and with a recorded consent statement from the speaker.

Where to use them

All three models run through Azure Speech in Microsoft Foundry, in public preview, and the voices are served from 14 Azure regions. The two voice models are also on OpenRouter, at the same prices. Microsoft lists Vercel, LiveKit (marked coming soon) and Azure Voice Live among the other routes to the models, and its MAI Playground runs a demo called Chatter on them.

A live transcriber joins Microsoft AI’s voice models

Microsoft AI already had speech models of its own: MAI-Voice-1 reached Microsoft Foundry alongside the new MAI-Transcribe-1 on 2 April 2026, and MAI-Voice-2 followed on 2 June. MAI-Transcribe-2-Streaming is its first model that listens live. “Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends,” the launch post says. The transcription side has an independent first place on Artificial Analysis’s accuracy ranking; the voice side’s quality figures are Microsoft’s own. For the batch model released a month earlier, see MAI-Transcribe-2; for the rival voice platform Microsoft compares Flash with, ElevenLabs.

Questions people ask

What is MAI-Transcribe-2-Streaming?
MAI-Transcribe-2-Streaming is Microsoft AI's live speech recognition model, released on 1 October 2026. It turns audio into text while the speaker is still talking, returning partial transcripts that update and final ones that confirm each segment, in 60 languages with automatic language detection. Artificial Analysis ranks it first of 38 streaming models for accuracy, at a 2.51% word error rate.
How much do the new MAI speech models cost?
MAI-Transcribe-2-Streaming costs $0.54 per hour of audio at an introductory price that runs to 31 December 2026, 5.4 times the $0.10 an hour, also an introductory price, of the batch model MAI-Transcribe-2. MAI-Voice-2.1 costs $22 per million characters of text and MAI-Voice-2.1-Flash $15, the same prices as MAI-Voice-2 and MAI-Voice-2-Flash.
What is the difference between MAI-Voice-2.1 and MAI-Voice-2.1-Flash?
MAI-Voice-2.1 is Microsoft AI's most expressive text-to-speech model, aimed at long-form work such as audiobooks and voice-overs, with about 550 milliseconds of model inference latency. MAI-Voice-2.1-Flash is the low-latency version for live conversation, such as call-centre agents and voice assistants, with about 45 milliseconds of model inference latency and streaming output that lets a caller interrupt it.
Can MAI-Voice-2.1 clone a voice?
Yes, under a gated process. Microsoft's documentation describes instant voice cloning from a 5 to 60 second reference clip, available only after approval through Azure's limited access review and with a recorded consent statement from the speaker.