Build8 publishers2 min readPublished Updated
Microsoft's MAI speech lineup documents real-time use only on the voice-output side
Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.
The Engineer · Build desk

What happened
- MAI-Voice-2 comes with broad language coverage and controls over prosody and emotion in the generated voice.
- According to a dev.to review, Microsoft's official documentation does not list MAI-Transcribe-2-Streaming, MAI-Voice-2.1 or MAI-Voice-2.1-Flash as product names.
- The MAI-Transcribe-2 catalog entry mentions limited-time pricing, but the review does not include the amounts or the terms.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Live captioning and agent-assist designs cannot assume a MAI-Transcribe-2 streaming endpoint yet; recorded-audio workflows are the ones that can adopt it on its current documentation.
- exposure Specs or vendor plans written against MAI-Transcribe-2-Streaming or MAI-Voice-2.1 depend on products Microsoft has not documented and need rewriting before a build starts.
- cost Teams cannot model MAI-Transcribe-2 unit costs from public figures, and the limited-time framing means any rate they are quoted now has an end date.
A live voice agent has two places where users feel delay. On the way in, the caller's speech has to become text while they are still talking. On the way out, the reply has to become audio before the pause gets awkward. Microsoft has a documented model for the outbound side: MAI-Voice-2-Flash, the lower-latency variant it intends for real-time voice-agent scenarios [6]. For the inbound side, a dev.to review of Microsoft's documentation found no native real-time streaming transcription advertised for MAI-Transcribe-2 [5]. Of the three documented models, the only one framed for real time is on the output side [1].
MAI-Transcribe-2 fits work where the audio already exists. The review says it may suit recorded audio, or workflows where transcription runs once an audio segment is available [10]. Diarization shows who said what. Word-level timestamps let a reviewer find the exact point in a long recording where a statement was made [12]. The review names post-call analysis, meeting notes, quality review and searchable audio archives as uses for those features [14].
I think putting transcription and voice response in separate models is the right design. The review treats them as separate implementation decisions with different latency and quality requirements [13]. A pipeline that reviews yesterday's support calls can wait for a batch job. A voice agent in the middle of a conversation cannot.
The obvious workaround for live use is to cut incoming audio into short segments, on a fixed interval or on silence, and submit each one as it closes. That is batch transcription on a timer. The caller-side delay can be no shorter than one segment plus one request. Whether speaker labels stay consistent across separately submitted segments needs measuring before anyone builds on it. The review's own advice is narrower: do not design a live transcription experience around an assumed MAI-Transcribe-2 streaming endpoint until Microsoft documents one [9].
The 60-language figure from the catalog entry is a list of supported languages [3]. Support and accuracy are separate claims. For the number to mean anything for a given contact centre, the model has to hold up on that centre's own calls. The review tells teams to test representative audio, languages, accents, noise conditions and conversation formats before relying on the output in customer-facing or operational processes [11].
What to watch
- A streaming endpoint for MAI-Transcribe-2 appearing in Microsoft's official documentation.
- Published amounts and an end date for the limited-time MAI-Transcribe-2 pricing.
- Independent per-language accuracy results for MAI-Transcribe-2 on noisy, multi-speaker call audio.