Invest1 distinct publisher3 min readPublished
The early-bird rate on MAI-Transcribe-2 sits 72% below the line's April price, which saves a 100,000-hour buyer $26,000 a year and takes 26 cents out of every 36 a standalone vendor charged for the same hour.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
An early-bird rate is a term rather than a price, and the whole question is which of the two 10 cents an hour turns out to be [1][2]. The cost side of the answer sits in throughput: Microsoft says the model clears audio 10 times faster than OpenAI's GPT-Transcribe, seven times faster than ElevenLabs' Scribe v2 and five times faster than Google's Gemini 3.5 Transcribe [14], and in batch work that ratio is the unit cost, because a model running at 300 times real time burns a fraction of the GPU-hours of one running at 30 [15]. Mustafa Suleyman told The Verge in April that the first model in the line came out of a 10-person team and ran at half the GPU cost of other state-of-the-art models [18]. If those multiples hold, a dime an hour can sit above marginal cost and stay there. If they do not, this is acquisition spend with a renewal date attached.
For anyone selling transcription by the hour, the number that matters is 36 divided by 10: a vendor that matched Microsoft's April rate now needs 3.6 times the volume to hold revenue flat [1]. Scale that down and it gets starker, since at 10 cents an hour a full year of continuously recorded audio, all 8,760 hours of it, bills at $876 [3], and a $1m revenue line requires 10 million hours of customer audio [4]. That is the arithmetic that makes standalone speech-to-text hard to fund, or rather, the more useful version is that it makes the features the funding, and Microsoft has moved speaker diarization, word-level timestamps, keyword biasing for domain jargon and automatic language identification into the base rate [6], which is the same menu the specialists have historically billed as extras.
The buyer, meanwhile, barely notices. On the 100,000 hours of call-center audio Microsoft uses as its own illustration, the bill goes from $36,000 to $10,000 [3], a saving of $26,000 a year [2], which is less than a single hire and nowhere near the cost of revalidating a compliance workflow. The volume moves on defaults, not on price, and Microsoft owns Teams and its meeting audio, Nuance and its clinical documentation, and the Azure speech services enterprises already call [17]. Bloomberg reported in July that MAI models had begun answering a portion of prompts in Word and Excel, products previously advertised as running on OpenAI and Anthropic [16], which is what substitution looks like when nobody signs anything.
What the evidence does not yet establish is whether the transcript is good enough for the specific buyer. The 5.2% average word error rate that Microsoft ranks first with comes from FLEURS [8], a benchmark built from native speakers reading about 2,000 sentences per language [9], and it is up from the 3.7% reported in June, roughly 41% higher in relative terms [10][5], with no per-language breakdown in the release [11]. Microsoft's other yardstick is Artificial Analysis, where it says the model now ranks second [12] on a leaderboard that had the previous version third at 2.4% [13]. A buyer whose traffic is Hinglish or Spanglish, the two cases Microsoft names for its code-switching mode [7], is buying an average.
The release names four rivals and omits Deepgram, AssemblyAI, Speechmatics and Rev [19], which reads to me as a judgment that those firms are not competing on API list price. The thesis fails if 10 cents expands the pool rather than dividing it: at $876 for a year of continuous audio, categories nobody bothered transcribing become worth transcribing, and 3.6 times the volume is a target the specialists can actually hit.
Ranked by verification strength, evidence, and original report placement.
Microsoft AI released MAI-Transcribe-2 on September 3, a speech-recognition model priced at 10 cents per hour of audio.
The 10 cents is an early-bird rate and undercuts the $0.36 an hour Microsoft charged for the first model in the line five months earlier by roughly 72%, according to VentureBeat's reporting on the launch.
For an enterprise running 100,000 hours of call-center audio a year, a modest volume for a large bank or telecom, the annual bill falls from $36,000 to $10,000.
The model transcribes 60 languages, up from 43 in June's MAI-Transcribe-1.5 and 25 in the April original, and Microsoft says it was built for background noise, low-quality recordings and overlapping speech.
It ships through Microsoft Foundry, the company's model marketplace, where the MAI multimodal family sits alongside more than 11,000 models from OpenAI, Anthropic, Meta, Google and others, and through MAI Playground.
Bundled into the base price are speaker diarization, word-level timestamps, keyword biasing for domain jargon such as drug names or product codes, and automatic language identification, features specialty vendors have historically charged extra for.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Meta folds diarization and endpointing into the transcript stream for $0.18 an hour2 distinct publishers
build
The demo-best voice engine finished last: 12,247 calls argue for buying on completion rate1 distinct publisher
leadership
Meta prices streaming transcription at a fifth of Google Cloud's standard rate1 distinct publisher
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One release, one relay
The checkable half of this story is solid: prices, dates, language counts, the feature bundle, the products MAI models have moved into. The accuracy and speed half rests entirely on Microsoft describing its own results, including its account of where it sits on someone else's leaderboard, and pivotnews.ai is the only newsroom here, citing VentureBeat, Bloomberg and The Verge rather than testing anything. The piece is candid about what is missing, which is unusual and worth crediting: no per-language error rates, no diarization error rate, nothing on streaming, no expiry on the launch price.
Shipping, buyers unnamed
Availability is essentially the whole of it. The model is live through Foundry and MAI Playground, and the only usage anywhere in this reporting is secondhand: Bloomberg's July account of MAI models handling some Word and Excel prompts. No buyer is named and no volume is disclosed. Teams meeting audio, Nuance's clinical documentation and Azure speech give Microsoft obvious places to route work, but owning the channel is not the same as traffic through it.
Read speech sold as call-center speech
Microsoft's framing is first on FLEURS and second on Artificial Analysis. The number underneath moved the wrong way, from 3.7% to 5.2%, and the test it leads on is native speakers reading sentences aloud, not the overlapping crosstalk the release says the model was built for. The price claim needs no discount: it is checkable and the buyer arithmetic holds. What is stretched is the distance between read-speech leadership and production audio, alongside a rate labelled early-bird with no expiry and no standard price behind it.
Seller sets the price and the scoreboard
Microsoft picked the rate, and it picked the benchmarks and the competitors it gets measured against too. The release measures itself against OpenAI, Google and ElevenLabs, skips the four specialists a 10-cent hour squeezes hardest, and quietly declines to claim it beats Alibaba on accuracy. Suleyman's half-the-GPU-cost line and the substitution of MAI models into Word and Excel point in one direction: cheap transcription is a lever against the partners Microsoft used to advertise as much as an offer to buyers.
Enough for the price, not the accuracy
One outlet, one launch, and an attribution chain that runs back to Microsoft for most of the load. We would act on the price, the language count, the feature bundle and the fact of the cut. We would not yet price conversational accuracy, diarization quality, streaming readiness, or the odds that 10 cents survives the launch window.