Build1 publisher2 min readPublished
LiveKit's Microsoft AI plugin brings MAI-Voice-2-Flash to voice agents that hold an Azure Speech key
LiveKit Agents now speaks in Microsoft's MAI voices, including MAI-Voice-2-Flash, through an official plugin that covers text-to-speech only. Teams no longer write their own adapter, but every deployment still needs an Azure Speech key and region.
The Engineer · Build desk

What happened
- The plugin installs as the livekit-agents[microsoft-ai]~=1.8 package extra for the LiveKit Agents framework.
- Microsoft's Foundry communications list LiveKit among the platforms where MAI models, including voice and speech models, are available.
- The documentation lists no price for MAI voice through LiveKit and no concurrency limits, language coverage, regional availability or latency targets.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability A team can now trial MAI voices against its current speech provider without leaving the LiveKit agent architecture it already runs.
- constraint Every deployment now depends on a working Azure Speech resource, and an agent without the key and region set has no voice at all.
- decision To get a per-call voice cost, a team has to price the Azure side and the LiveKit side separately before it commits.
According to a dev.to write-up of LiveKit's documentation, the plugin goes into a LiveKit agent session as the microsoft_ai.TTS provider [4]. The session sends the text its agent generates to Microsoft AI TTS and gets synthesized audio back for the conversation [13].
The plugin reads two environment variables, MICROSOFT_AI_TTS_API_KEY and MICROSOFT_AI_TTS_REGION [7]. The variable names say Microsoft AI, but the values are an Azure Speech resource key and an Azure region [5]. A team with no Azure footprint has to provision a Speech resource before its agent can speak [5].
The scope is speech output. The write-up describes a direct TTS integration [2]. Its suggested uses include automated phone menus [18].
The documented example shows one model, one voice and one sample rate [11]. For that example to carry over to a real deployment, the caller's language needs a MAI voice, the team's Azure region has to serve the model, and the audio path has to work at the example's 24 kHz or the plugin has to accept another rate [3][11]. The write-up recommends testing the chosen voice under realistic call or session volumes before committing a customer-facing workflow [17].
LiveKit says its Inference system supports pay-as-you-go pricing across providers and global concurrency management [9]. The plugin path in the documented example still authenticates with the team's own Azure key [5].
I think this is the right tradeoff for a team already running LiveKit agents. The integration code moves behind a documented provider interface, and what remains is account setup and testing [12][17]. The cost is a second vendor account in the deployment path [5].
What to watch
- LiveKit plugin documentation that lists the full set of MAI voices, languages and supported sample rates.
- A published price for MAI voice usage through LiveKit, or a statement on whether MAI voice is sold through LiveKit Inference.
- Latency and concurrency figures for MAI-Voice-2-Flash measured at real call volumes.