Microsoft AI has launched its first streaming MAI transcription model together with MAI-Voice-2.1 and the faster MAI-Voice-2.1-Flash. The three-model release gives developers an in-house Microsoft stack for incremental speech recognition, expressive multilingual synthesis, and latency-sensitive voice agents.
On October 1, 2026, Microsoft AI released MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. MAI-Transcribe-2-Streaming is the company's first streaming transcription model: it accepts continuous audio over WebSocket, returns changing partial transcripts while a person is speaking, and confirms final segments without making an application wait for the whole utterance. Microsoft says the model supports 60 languages with continuous automatic language detection.
Microsoft positions the streaming model around the accuracy-latency trade-off measured by Artificial Analysis. Unite.AI's account of the launch cites the September 28 leaderboard at 2.5% final word error rate, 2.8% for the first partial, and 0.13 seconds to final transcription, while Microsoft says initial hypotheses can arrive in just over 100 milliseconds. The independent benchmark combines AA-AgentTalk, VoxPopuli, and Earnings22; production results will still depend on audio conditions, network delay, and the rest of an application's stack.
For speech generation, MAI-Voice-2.1 is Microsoft's higher-fidelity, expressive tier for long-form narration and brand audio, while MAI-Voice-2.1-Flash targets high-volume and interactive workloads. Both cover 23 languages and 26 locales, maintain one speaker identity across language switches, and offer gated voice cloning from a short reference sample. Microsoft reports that Flash can generate up to 45 seconds of audio with 150 milliseconds of end-to-end latency, with pricing listed at $22 per million characters for Voice 2.1 and $15 for Flash.
The models are available through Microsoft Foundry and the MAI Playground, with additional access through services including Vercel, OpenRouter, and Azure Voice Live. Microsoft also built a Chatter demo that combines streaming recognition and synthesized speech into a live assistant. Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash gives developers Microsoft-built components on both sides of a voice-agent loop: listening before a speaker finishes, then answering with a low-latency generated voice.