Microsoft launches MAI-Transcribe-2-Streaming, a real-time speech-to-text model that tops the accuracy charts
First streaming model with leading accuracy
Microsoft released MAI-Transcribe-2-Streaming on the first of October, its first continuous-streaming speech-to-text (STT) model. It entered directly at number one out of 38 contenders on Artificial Analysis's AA-WER Streaming benchmark, leading on both final-transcript accuracy and first-partial-transcript accuracy. The model targets voice agents, live captioning and dictation — scenarios where latency defines the user experience.
How it works under the hood
It is the real-time twin of the batch MAI-Transcribe-2 released in September. The model supports 60 languages with continuous automatic language detection. Audio flows in continuously; text flows out while the speaker is still talking. The first hypotheses, called partials, are emitted a little over 100 milliseconds after audio arrives and update as more context arrives until a stable final transcript is settled. That lets an agent start reasoning or invoking tools mid-sentence. According to Microsoft's internal testing, words appear at twice the rate of the closest competitor.
What the benchmark actually shows
The AA-WER Streaming benchmark draws on roughly 8 hours of audio across three slices: fifty percent AA-AgentTalk, twenty-five percent VoxPopuli, twenty-five percent Earnings22. Latency is measured from the end of speech as detected by SileroVAD. The result: