Microsoft AI
MAI-Transcribe-2
speech recognition at ten cents an hour

Key facts
- 3 Sep 2026Microsoft AI
- Published
- $0.10per hour of audio, limited time
- Price
- 5.2%WER, 60 languages
- FLEURS
- Firston FLEURS, Microsoft's claim
- Position
Microsoft AI's second in-house transcription model prices speech recognition at ten cents an hour and claims first place on FLEURS across 60 languages at a 5.2% word error rate. The claim of fastest, most accurate and cheapest is Microsoft's own; the pricing is real and the pressure it puts on transcription incumbents is the story.
What it is
MAI-Transcribe-2 is Microsoft AI’s speech recognition model, published on 3 September 2026 under a headline that leaves no room for modesty: the fastest, most accurate and cheapest speech recognition model in the world. The claims are Microsoft’s own and deserve reading as such, but the two numbers underneath them are concrete. On FLEURS, the multilingual benchmark that tests transcription across 60 languages, Microsoft reports first place at a 5.2% word error rate. And the price is $0.10 per hour of audio on a limited-time rate.
Why the price is the story
Ten cents an hour reframes what transcription costs. At that rate an entire working day of meetings transcribes for less than a dollar, a full podcast back catalogue for a few tens of dollars, and the economics of building transcription into a product change shape: the cost stops being a line item and becomes a rounding error. The pressure lands on the incumbent transcription APIs and on the middle of the market, the services that priced per minute in an era when accurate multilingual speech recognition was scarce. The limited-time label on the rate is worth watching: introductory pricing that doubles later has become a pattern across the model market this year.
Where it fits
Microsoft AI has been shipping in-house models at a steady clip under Mustafa Suleyman, across image generation and now speech, alongside the OpenAI models Microsoft resells through Azure. A transcription model of its own, priced aggressively, gives Microsoft a wedge into every workflow that starts with recorded speech: meetings, calls, media, accessibility. For the wider field, the direction counts for more than the benchmark claim: multilingual transcription at commodity prices is now table stakes, and the competition moves to what a product does with the transcript. Our audio models hub tracks the rest of the field.