Microsoft expanded its MAI family with three voice models built for agents that listen and reply without long pauses. The central release is MAI-Transcribe-2-Streaming, priced at 54 cents per audio hour and producing first transcript hypotheses within 320 milliseconds on average. The model supports more than 60 languages with automatic language detection. For business, this matters because live transcription lets an agent show captions or start work before a caller finishes speaking.
Streaming transcription plus two speech models
MAI-Transcribe-2-Streaming accepts speech through a WebSocket and updates the transcript continuously while audio keeps arriving. When the speaker stops, the model marks the transcript as final, which suits live captions and early request processing. Microsoft lists the model on Vercel AI Gateway, where pricing and language coverage are published. The company cautions that actual response speed also depends on network connection and the reply-generation system. That caveat matters for contact centers and apps where latency varies by infrastructure.
The transcription pipeline differs from batch processing because it returns provisional results during speech rather than waiting for silence. Last month Microsoft introduced MAI-Transcribe-2 at 10 cents per audio hour, more than five times cheaper than the streaming version. The price gap reflects extra computation for continuous updates while audio is still incoming. The older approach starts only after the utterance ends, which saves cost but adds waiting time. Developers therefore choose between lower transcription cost and faster turn-taking.
Two companion releases cover the output side of a voice agent. MAI-Voice-2.1 targets more expressive and higher-fidelity speech, while MAI-Voice-2.1-Flash trades some expressiveness for faster response and lower cost. Vercel AI Gateway lists MAI-Voice-2.1 at $22 per million characters and Flash at $15 per million characters. Both text-to-speech models support 23 languages. Microsoft positions the trio with its Mai-Thinking-1 reasoning model, which reads transcripts and decides what the agent should do next.
What this means for companies building voice agents
For firms deploying support lines, booking assistants, or in-app voice controls, the stack separates three controls: recognition quality, reasoning, and voice output. A team can keep transcription streaming for responsiveness while selecting Flash voice to contain character costs, or choose higher-fidelity voice for premium customer-facing roles. Language coverage also shapes rollout: 60-plus languages for input versus 23 for output means some markets get understanding before natural replies. Small pilots can start with one language and one voice, while larger operations can mix models by use case.
Cost and latency decisions need testing against real traffic rather than list prices alone. Streaming transcription costs 54 cents per audio hour against 10 cents for batch, so high-volume recording without a need for instant action may stay on batch. Voice output is billed per million characters, which makes long confirmations, disclaimers, and repeated prompts more expensive than short answers. Buyers should ask vendors how WebSocket stability, automatic language detection errors, and downstream model delays affect the 320-millisecond hypothesis figure. The news by itself does not guarantee humanlike conversation in every network environment.
A practical marker will be whether Microsoft moves Copilot agents in Excel, Outlook, and related products onto MAI models. Chief executive of Microsoft AI Mustafa Suleyman has linked MAI development to reducing payments to Anthropic and OpenAI, despite Microsoft investments in both firms. A shift of flagship assistants to in-house transcription, reasoning, and voice would signal confidence in quality and unit economics. Until then, outside developers provide the early test of reliability at scale.
