Microsoft is taking another swing at the speech-to-text market with MAI-Transcribe-2, a new transcription model designed to combine high accuracy with much faster processing and lower costs.
The model, announced by Microsoft AI on September 3, is positioned as the company’s most capable transcription model yet. Microsoft says it can handle everything from clinical notes and legal documentation to accessibility tools and closed captions, while supporting 60 languages and a wide range of real-world audio conditions.
Microsoft targets speed, accuracy and cost
MAI-Transcribe-2 is built around a fairly straightforward pitch: developers shouldn’t have to choose between transcription quality, processing speed and price.
Microsoft says the model ranks first on the FLEURS multilingual benchmark across 60 languages, achieving an average Word Error Rate of 5.2%. It also ranks second on Artificial Analysis’ Word-Error Rate leaderboard and sits on the benchmark’s accuracy-latency Pareto frontier.
The company claims MAI-Transcribe-2 can process audio up to 10 times faster than leading competitors. In evaluations conducted by Artificial Analysis, Microsoft says it was 10 times faster than OpenAI‘s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2 and five times faster than Gemini 3.5 Transcribe, while delivering higher accuracy in those comparisons.
That speed could matter particularly for applications handling long recordings, where transcription latency can quickly become a bottleneck.
Built for messy, real-world audio
The model also goes beyond basic speech recognition. MAI-Transcribe-2 includes speaker diarization, allowing it to distinguish between speakers and associate words with the appropriate person. It also supports word-level timestamps, which can make it easier to search, edit and synchronize transcripts with audio or video.
Microsoft has added keyword biasing for specialized terminology, names, abbreviations and other words that can be difficult for a general transcription model to recognize correctly from context.
There are also configurable transcription styles. The “verbatim” mode preserves filler words and false starts, which can be useful for compliance and analysis, while the “clean” mode removes those elements to produce more readable captions, notes and transcripts.
MAI-Transcribe-2 can also automatically identify the language being spoken and handle code-switching. Microsoft specifically points to mixed-language conversations such as Hinglish and Spanglish, making the model potentially useful for multilingual environments where speakers naturally switch languages during a conversation.
Microsoft says the model is designed to remain robust in noisy environments rather than being limited to controlled recording conditions.
$0.10 per hour at launch
Perhaps the most aggressive part of Microsoft’s announcement is the pricing.
MAI-Transcribe-2 launches at $0.10 per hour of audio as a limited-time offer through the end of 2026. Microsoft says its efficiency allows it to offer what it considers the most competitive price on the market while maintaining its claimed accuracy and throughput advantages.
For developers processing large volumes of audio, that combination could make the model particularly interesting. A lower per-hour cost becomes much more meaningful when transcription is being performed across thousands or millions of minutes of recordings.
The model is available to try through Microsoft Foundry, Microsoft’s MAI Playground and OpenRouter.
MAI-Transcribe-2 also fits into Microsoft’s broader push to build its own family of MAI models rather than relying entirely on third-party AI systems. With transcription increasingly becoming a basic layer for meeting assistants, accessibility tools, media applications and AI agents, Microsoft’s focus on speed and cost suggests it is targeting the infrastructure underneath those experiences as much as the end-user applications themselves.
For now, the biggest claims around MAI-Transcribe-2 come directly from Microsoft’s own evaluations, so independent testing will be important in determining how the model performs against competitors across different accents, languages, recording conditions and specialized terminology. But on paper, Microsoft is making a clear case: transcription doesn’t have to be expensive or slow to be highly accurate.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
