Microsoft's New Goal: Voice AI Agents That Speak Like Humans

· AI · Cem Koyluoglu

Microsoft Corp. Expands MAI Model Family by Introducing First Live Transcription Model Alongside Two Other Text‑to‑Speech Models.

TL;DR Microsoft has expanded its MAI AI model family by adding a live transcription model, enabling developers to build voice agents that can listen and respond in real time. The new MAI‑Transcribe‑2‑Streaming model supports over 60 languages, offers low latency, and is priced at $0.54 per hour of audio. This move signals Microsoft’s push to reduce reliance on third‑party providers like OpenAI and Anthropic.

Key Highlights • MAI‑Transcribe‑2‑Streaming receives spoken input over WebSocket, updates transcription continuously, and delivers final transcript after speaker finishes. • The streaming model costs $0.54 per hour, supports 60+ languages, and returns initial hypotheses in ~320 ms, while the non‑streaming MAI‑Transcribe‑2 is $0.10 per hour. • Microsoft’s MAI family also includes two text‑to‑speech models—MAI‑Voice‑2.1 ($22/million chars) and MAI‑Voice‑2.1‑Flash ($15/million chars)—supporting 23 languages. • The rollout reflects Microsoft’s strategic shift to prioritize its own MAI models over external providers, aiming to power Copilot agents in Excel and Outlook.


Microsoft expands MAI model family with live transcription capability

Microsoft Corp. announced the addition of its first live transcription model to the MAI suite, joining two existing text‑to‑speech models. The new lineup is aimed at developers building voice agents that can listen to human speech and respond instantly, mimicking natural conversation.

Live transcription model

The flagship of the three models is MAI‑Transcribe‑2‑Streaming. Microsoft explained that the model receives spoken input over a WebSocket connection and continuously updates a transcription as the speaker talks. When the speaker finishes, the final transcript is delivered. This enables applications that provide live captions or begin processing a request before the user has completed their utterance.

Pricing and performance

MAI‑Transcribe‑2‑Streaming is listed on Microsoft’s Vercel AI Gateway platform at $0.54 per hour of audio. It supports more than 60 languages and can automatically detect the spoken language. Initial transcription hypotheses are typically returned within an average of 320 milliseconds, though Microsoft cautioned that actual latency depends on network conditions and the underlying AI system.

Earlier transcription offering

Microsoft is not new to speech transcription. Last month the company introduced MAI‑Transcribe‑2 at $0.10 per hour, roughly five times cheaper than the streaming version. The cost difference stems from the streaming model’s need to generate interim results continuously while the audio stream is ongoing, whereas the standard model waits for the speaker to finish before processing.

Text‑to‑speech models

Beyond transcription, the MAI family includes two text‑to‑speech models. MAI‑Voice‑2.1 is positioned as the higher‑quality option, while MAI‑Voice‑2.1‑Flash trades a small amount of fidelity for faster response times and lower cost. According to Vercel AI Gateway listings, MAI‑Voice‑2.1 is priced at $22 per million characters and the Flash variant at $15 per million characters. Both models support 23 languages.

Strategic shift away from third‑party providers

The rollout highlights Microsoft’s accelerating move away from external model providers such as OpenAI Group PBC and Anthropic PBC. In July, Microsoft AI Executive Vice President Mustafa Suleyman expressed growing concern over the expenses associated with OpenAI and Anthropic’s flagship models. He directed Microsoft’s AI researchers to prioritize the MAI model family, with the goal of powering Copilot agents in products like Excel and Outlook. In an interview with Bloomberg, Suleyman said, “We are paying a lot to Anthropic, so our aim is to reduce that cost and eventually eliminate it entirely.”

Architecture of voice agents

The new models give developers granular control over the quality, latency, and cost of voice agents. MAI‑Transcribe‑2‑Streaming addresses the first step—recognizing and understanding spoken input. The MAI‑Voice models handle the final step—producing an audible response. The intermediate decision‑making step is performed by a standard large language model, Microsoft’s MAI‑Thinking‑1, which reads the transcription and determines the appropriate action. By combining three distinct models, developers can fine‑tune each stage of the voice interaction pipeline.

Source: https://siliconangle.com/2026/10/01/microsoft-targets-ultra-realistic-voice-agents-with-its-first-streaming-transcription-model/

← All tech news