By Mike Wheatley
Publication Date: 2026-10-02 01:31:00
Microsoft Corp. today expanded its MAI artificial intelligence model family with its first streaming transcription model, debuting alongside two others focused on text-to-speech.
They’re designed for developers who want to build voice agents that can listen to people’s voices and reply instantly, similar to how humans talk to one another.
The most consequential of the three is MAI-Transcribe-2-Streaming, which the company said accepts human speech via a WebSocket, transforming it into a transcript that’s continuously updated as the person keeps on talking. Once the person has finished speaking, it will confirm that the transcript is final. In this way, the model can power applications that display live captions or start processing a user’s request before they’ve finished saying it, Microsoft said.
MAI-Transcribe-2-Streaming is listed on Microsoft’s Vercel AI Gateway and priced at 54 cents per audio hour, the company said. It supports more than 60 languages, and can…


