Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
Oct 2, 2026, 10:09 PM · MarkTechPost

Microsoft’s first streaming STT model tops Artificial Analysis for accuracy and partials — faster words for voice agents, at a premium to xAI and Meta.
Why it matters
Michal Sutter at MarkTechPost reports that Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash text-to-speech. It’s the real-time sibling of batch MAI-Transcribe-2 from September: 60 languages with continuous auto-detect, audio in and text out while the speaker is still talking.
Artificial Analysis ranks it #1 of 38 on AA-WER Streaming for both final and first partial transcripts — 2.5% WER at 0.13s to final and 0.12s to first partial on a mix of AA-AgentTalk, VoxPopuli and Earnings22. Runners-up include Grok Voice Transcribe 2.0 (2.7% at 0.49s) and Muse Voice Transcribe (3.1% at 0.16s). Cartesia Ink-2 is faster to final (0.07s) at 4.0% WER.
Introductory price is $0.54 per hour of audio through end of 2026 — higher than xAI and Meta on the comparison table, roughly in line with Google’s estimated rate. Integration paths: OpenAI-compatible Realtime WebSocket and Azure Speech SDK. Public preview, no SLA, no open weights.
From the desk
We’re for streaming STT that makes partials trustworthy. Voice agents don’t wait for the period; they start tool calls mid-sentence. Matching 2.5% WER on first partial and final is the headline that matters more than a pure speed crown Cartesia can still claim.
Useful AI here is accessibility and product fluency: live captions, dictation, hands-free agents. Pairing with MAI-Voice-2.1-Flash (45s of audio at 150ms end-to-end for $15 per 1M characters in the write-up) sketches a full Microsoft voice loop. That’s competitive pressure on everyone selling realtime voice stacks.
The downside if closed streaming STT wins is the usual: no weights, no SLA in preview, and pricing that sits above Meta and xAI while accuracy leads. Enterprises will pay for WER; startups may not. Continuous language detection across 60 languages is powerful and raises the familiar surveillance/recording questions any always-on mic product inherits.
I’m watching whether Microsoft keeps the #1 AA rank once rivals retrain, and whether LiveKit support and an SLA arrive before buyers treat this as production-critical.
Context
Microsoft says internal tests show words appearing 2x faster than its closest competitor. The model sits on Artificial Analysis’s accuracy-versus-latency Pareto frontier in the MarkTechPost summary.
Who feels it
- Voice-agent builders
- A top-ranked streaming STT with Realtime API familiarity — preview caveats and higher $/hour than some rivals.
- xAI, Meta, Google voice stacks
- Accuracy bar just moved; latency and price remain competitive levers.
- Accessibility & captioning products
- Stronger partials mean captions that update sooner without as much garbage text.
What to watch
- AA-WER Streaming rankings after the next rival releases
- Move from public preview to SLA-backed GA
- LiveKit support and diarization features relative to Muse’s multi-speaker claims
Companies: Microsoft