Product1 distinct publisher3 min readPublished
Muse Voice Transcribe claims diarization, endpointing and mid-sentence language switching in a single pass at $3 per 1,000 audio minutes, which makes the buying question less about accuracy scores than about which 25 languages Meta validated.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The word doing the work in Mark Zuckerberg's post is "natively": diarization and endpointing in a single model [2]. If speaker labels, segment boundaries and language identification all come out of the same pass, they cannot contradict each other, and the contradiction is where the cleanup time goes when those jobs are wired together from separate parts. The tradeoff arrives in the same breath. One model means one failure surface, and you cannot replace the speaker labeling without replacing the transcription underneath it.
The price makes that a real decision rather than a thought experiment. At $3 per 1,000 audio minutes [6], you are paying $0.003 a minute, or about $0.18 an hour of audio [1]. The hour-long, 20-plus-speaker session Meta describes [4] costs roughly eighteen cents to run through the API [4]. A team pushing 40 hours of recordings a week is at 2,080 hours a year, or about $374 [5]. At that level the scarce resource is not budget, it is the person who checks whether the labels are right before anyone quotes the transcript back at a customer.
What the material does not contain is any independent measurement. The state-of-the-art claim for streaming speech-to-text is Zuckerberg's own [2], the adaptive delay description is his [5], and the third-party account is Engadget's write-up of a demo video Zuckerberg posted [1]. No word error rates, nothing per language, nothing on diarization when two people talk over each other. Meta says the model was trained across more than 70 languages with 25 validated at launch [3], which leaves at least 45 trained but unvalidated [2] and puts validated coverage under 36 percent of the trained set [3]. "Validated" is quietly the most important word in the release.
Distribution tells you who this is actually for. Muse Voice Transcribe is live in Meta's desktop app dictation and Muse Code, and available through the Meta Model API [11][6], with a demo on the research blog [8], while Google is putting Gemini 3.5 Transcribe into Android and eventually Chrome [9]. Meta's buyer is a developer who already runs a transcription pipeline and can point it somewhere else, or a Mac user whose dictation gets better because the app powers voice features in other apps [7]. Google's buyer does not choose at all; the model shows up in the operating system.
The forcing function for a Monday rollout has two axes: whether your audio sits inside the 25 validated languages [3], and whether a human reads the transcript before anything downstream acts on it. Validated plus human review is the pilot case, and the pilot should be your own recordings run through both your current setup and this one. Validated plus automated downstream is a hold, because a single-model error has no second opinion to disagree with it. Outside the 25 with a human reading, it is usable if you accept doing the language validation Meta has not published. Outside the 25 and automated, not yet. The number worth collecting from that pilot is minutes a person spends fixing speaker attribution per hour of transcript, before and after. Word error rate on clean audio will flatter everyone.
Ranked by verification strength, evidence, and original report placement.
Meta has introduced Muse Voice Transcribe, its first real-time audio model, which Meta says can handle dictation and transcription for more than 20 speakers and multiple languages at once; Engadget reported that in a video Zuckerberg shared, the transcription automatically distinguishes between speakers, switches languages and picks up code-switching.
In a post dated September 1, 2026, Mark Zuckerberg described Muse Voice Transcribe as Meta Superintelligence Lab's first real-time audio perception model, rolling out that day, 'SOTA in streaming speech-to-text', handling speaker diarization and endpointing natively in a single model.
Zuckerberg said the model was trained across 70+ languages, with 25 validated at launch.
Zuckerberg said the model handles mid-sentence code-switching and manages hour-long sessions with 20+ speakers.
Zuckerberg said 'The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy.'
Muse Voice Transcribe is available to developers within Muse Code and Meta's Model API, priced at $3 for 1,000 audio minutes.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Meta budgeted up to $10bn a year with Anthropic even as Zuckerberg warns against rival AI labs1 distinct publisher
product
Meta's 2026 Capex Plan More Than Doubles Two Years of Spending; the Product It Funds Is Still on Paper1 distinct publisher
invest
Meta's coding agent has two prices: pay 18x more, or let it train on your repository1 distinct publisher
leadership
Meta prices streaming transcription at a fifth of Google Cloud's standard rate1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Two company posts, printed twice
Everything that carries weight here — the single-model diarization, the 70-plus training languages, the $3 price — comes from Zuckerberg's and Wang's launch-day posts and a video Meta chose, relayed by Engadget. Our coverage holds that same write-up twice, which doubles the paperwork and adds no verification. What would make the announcement checkable is entirely absent: no error rates, no latency, not even the names of the 25 validated languages.
Shipping, entirely inside Meta
This is not a preview. Dictation in the Meta AI Mac app, use inside Muse Code and a live Model API endpoint all exist on day one, which is more than most launch announcements can show. What is missing is anyone outside Meta: no third-party deployment, no usage figure, and Engadget's own caveat that flagship integration remains unannounced while Google routes its rival model through Android and Chrome.
State of the art, asserted
The shipped facts are solid — price, endpoints, surfaces, all checkable within a day. The superlatives ride on top of them without support: 'SOTA in streaming speech-to-text' arrives with nothing to compare it against, and the language claim reads impressively until you do the subtraction and find fewer than 36 percent of the trained languages were validated. Adaptive delay is described in prose and never in milliseconds.
Launch-day megaphone
Both original voices are selling. Zuckerberg broke a three-year silence on X to post the demo, Wang followed with a thread the same afternoon, and the timing places the announcement less than a week behind Google's comparable release. Engadget adds the useful caveat about flagship integration but runs no tests of its own, so the framing that reaches a reader is the one Meta wrote.
Certain what was said, not how it performs
We are on firm ground about what Meta announced, where it runs and what it charges, and the cost arithmetic derived from that price is ours and holds. Past that the ground gives way: one newsroom, one duplicated story, two promotional posts, and no way from here to know how the model behaves on messy audio in a language Meta has not named.