Microsoft AI
MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash
live transcription that tops the streaming accuracy board, and two new voices

Key facts
- 1 Oct 2026Microsoft AI, public preview
- Released
- 2.51% WERfirst of 38, Artificial Analysis
- Streaming accuracy
- $0.54 / hourintroductory, to 31 Dec 2026
- Transcription price
- $22 / $15per M characters, 2.1 / 2.1-Flash
- Voice prices
- 23one voice keeps its identity
- Voice languages
- 60with automatic detection
- Transcription languages
Microsoft AI released three speech models on 1 October 2026 for building voice agents that listen and talk in real time. MAI-Transcribe-2-Streaming turns live audio into text as people speak, in 60 languages, and ranks first for accuracy on Artificial Analysis's streaming board at a 2.51% word error rate. MAI-Voice-2.1 and MAI-Voice-2.1-Flash turn text into speech in 23 languages, at $22 and $15 per million characters, with Flash built for fast replies in live conversation.
Microsoft AI released three speech models on 1 October 2026, all aimed at voice agents: software that listens to a person, decides what to do and answers out loud, fast enough to feel like a conversation. MAI-Transcribe-2-Streaming turns live speech into text as it is spoken. MAI-Voice-2.1 and MAI-Voice-2.1-Flash turn text into speech, in 23 languages.
“A voice agent is a loop. It has to hear, understand, decide, and speak,” the launch post says. “And do it all within the window where a human still experiences the interaction as a conversation.” All three run through Azure Speech in Microsoft Foundry, in public preview.
MAI-Transcribe-2-Streaming writes while you talk
MAI-Transcribe-2-Streaming transcribes audio as a continuous stream, returning partial transcripts that update while the speaker talks and final transcripts that confirm each segment, in 60 languages with automatic language detection. It is the streaming sibling of MAI-Transcribe-2, the batch model Microsoft AI released on 3 September 2026, and its first partials arrive “in just over 100ms of receiving audio”, Microsoft says.
Developers connect through the Azure Speech SDK or over a WebSocket protocol compatible with OpenAI’s Realtime API. On the WebSocket route they send 16-bit mono audio in small chunks, and each session can last up to an hour, according to Microsoft Learn. Word-level timestamps, keyword biasing and speaker labels stay with the batch model, Microsoft’s comparison table shows.
It ranks first for accuracy on the streaming board
Artificial Analysis ranks MAI-Transcribe-2-Streaming first of 38 streaming models for accuracy, at a 2.51% word error rate on final transcripts and 2.52% on first partials, on its board read on 2 October 2026. The index mixes three test sets, around eight hours of voice-agent speech, European Parliament proceedings and company earnings calls, according to its methodology. Artificial Analysis reported it “taking the #1 spot” on 1 October.
| Model, rank on accuracy | Word error rate, final | Time to final transcript | Price per 1,000 minutes |
|---|---|---|---|
| MAI-Transcribe-2-Streaming, 1st | 2.51% | 0.128 s | $9.00 |
| Grok Voice Transcribe 2.0, 2nd | 2.73% | 0.490 s | $3.33 |
| Muse Voice Transcribe, 3rd | 3.06% | 0.163 s | $3.00 |
| Cartesia Ink Preview (external endpoints), 4th | 3.11% | 0.105 s | $4.00 |
| ElevenLabs Scribe v2 Realtime, 7th | 3.59% | 0.141 s | $6.50 |
| GPT Live Transcribe, 9th | 3.92% | 0.812 s | $17.00 |

Accuracy is the claim it wins. On the same board ten models return a final transcript sooner, led by Deepgram Flux at 0.021 seconds, and its $9.00 per 1,000 minutes is one of the higher prices: 28 of the 38 models cost less. Microsoft says it sits “on the Pareto frontier” for accuracy against latency.
Transcription costs $0.54 an hour
MAI-Transcribe-2-Streaming costs $0.54 per hour of audio, an introductory price that runs to 31 December 2026, according to its model card. That is 5.4 times the $0.10 an hour, also an introductory price, that Microsoft charges for the batch MAI-Transcribe-2, which waits for a recording to end before it transcribes.
| Model | What it does | Price |
|---|---|---|
| MAI-Transcribe-2-Streaming | Live speech to text | $0.54 per hour of audio |
| MAI-Transcribe-2 | Recorded speech to text | $0.10 per hour of audio |
| MAI-Voice-2.1 | Text to speech, most expressive | $22 per M characters |
| MAI-Voice-2.1-Flash | Text to speech, low latency | $15 per M characters |
At those rates, an hour of a voice agent listening costs $0.54 to transcribe, and 100,000 characters of spoken reply cost $1.50 on Flash.
Two voices that keep their identity across 23 languages
MAI-Voice-2.1 and MAI-Voice-2.1-Flash speak 23 languages, and one voice keeps the same identity in each of them with a native accent, Microsoft says. “Just ask it to speak English… then Mandarin… then German… and the speaker stays unmistakably the same,” the launch post says. Microsoft Learn lists 97 prebuilt voices, 50 male and 47 female, all usable with both models, and styles such as joy, excitement and empathy set through speech markup.
| MAI-Voice-2.1 | MAI-Voice-2.1-Flash | |
|---|---|---|
| Built for | Audiobooks, voice-overs, long-form speech | Call centres, voice assistants, live dialogue |
| Model inference latency, Microsoft | About 550 ms | About 45 ms |
| Price per million characters | $22 | $15 |
Flash streams its output, so a caller can interrupt it mid-sentence. The quality figures are Microsoft’s own. In blind side-by-side comparisons over 5,032 judgments, listeners preferred MAI-Voice-2.1 with new prompts 58.9% of the time against MAI-Voice-2 with old ones. In a test with 4,000 listeners, 50.3% rated the MAI voices as human-like as or more human-like than human recordings, with 2.1 and 2.1-Flash results combined, the model page says. Flash delivers “55% faster model inference” and is about 60% cheaper than comparable models, Microsoft adds, in a comparison against ElevenLabs’ published model documentation.

Voice cloning is gated. Microsoft’s documentation describes “instant voice cloning” from a 5 to 60 second reference clip, available only after approval through Azure’s limited access review and with a recorded consent statement from the speaker.
Where to use them
All three models run through Azure Speech in Microsoft Foundry, in public preview, and the voices are served from 14 Azure regions. The two voice models are also on OpenRouter, at the same prices. Microsoft lists Vercel, LiveKit (marked coming soon) and Azure Voice Live among the other routes to the models, and its MAI Playground runs a demo called Chatter on them.
A live transcriber joins Microsoft AI’s voice models
Microsoft AI already had speech models of its own: MAI-Voice-1 reached Microsoft Foundry alongside the new MAI-Transcribe-1 on 2 April 2026, and MAI-Voice-2 followed on 2 June. MAI-Transcribe-2-Streaming is its first model that listens live. “Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends,” the launch post says. The transcription side has an independent first place on Artificial Analysis’s accuracy ranking; the voice side’s quality figures are Microsoft’s own. For the batch model released a month earlier, see MAI-Transcribe-2; for the rival voice platform Microsoft compares Flash with, ElevenLabs.
Questions people ask
- What is MAI-Transcribe-2-Streaming?
- MAI-Transcribe-2-Streaming is Microsoft AI's live speech recognition model, released on 1 October 2026. It turns audio into text while the speaker is still talking, returning partial transcripts that update and final ones that confirm each segment, in 60 languages with automatic language detection. Artificial Analysis ranks it first of 38 streaming models for accuracy, at a 2.51% word error rate.
- How much do the new MAI speech models cost?
- MAI-Transcribe-2-Streaming costs $0.54 per hour of audio at an introductory price that runs to 31 December 2026, 5.4 times the $0.10 an hour, also an introductory price, of the batch model MAI-Transcribe-2. MAI-Voice-2.1 costs $22 per million characters of text and MAI-Voice-2.1-Flash $15, the same prices as MAI-Voice-2 and MAI-Voice-2-Flash.
- What is the difference between MAI-Voice-2.1 and MAI-Voice-2.1-Flash?
- MAI-Voice-2.1 is Microsoft AI's most expressive text-to-speech model, aimed at long-form work such as audiobooks and voice-overs, with about 550 milliseconds of model inference latency. MAI-Voice-2.1-Flash is the low-latency version for live conversation, such as call-centre agents and voice assistants, with about 45 milliseconds of model inference latency and streaming output that lets a caller interrupt it.
- Can MAI-Voice-2.1 clone a voice?
- Yes, under a gated process. Microsoft's documentation describes instant voice cloning from a 5 to 60 second reference clip, available only after approval through Azure's limited access review and with a recorded consent statement from the speaker.
More in Audio and Voice
All Audio →- Microsoft AIMAI-Transcribe-2speech recognition at ten cents an hour
- ElevenLabsElevenLabsthe voice infrastructure standard
- Alibaba QwenQwen-Audio-3.1Alibaba's upgraded audio stack, announced 23 September 2026
- Open and real-time audioStable Audio 3, GPT-Live, Nemotron ASR and Transcribe Arabic
- AlibabaHappyShrimp 1.0the name is a eulogy
- SunoSuno v6a flagship, an exploratory variant and a free mini, released together