Leadership1 distinct publisher3 min readPublished
Muse Voice Transcribe's benchmark lead rests on three-tenths of a point across about eight hours of English audio, which is thinner evidence than the $0.18 hourly rate sitting underneath it.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
The ratio is the part of this a buyer can plan around. Google Cloud's standard $0.96 an hour is 5.3 times Meta's $0.18 [15][4]. At 1,000 hours of audio a month the gap is $780, or $9,360 a year [16]. That is small enough that most teams have never had the line reviewed and large enough that a finance partner who has seen a leaderboard will ask why it exists. The harder part of the conversation is that "our vendor is more accurate" no longer answers it: Google's Gemini 3.5 Transcribe Live sits at 4% word error against Muse's 3.1%, about 22% more errors on the same test [3][2][17].
That test carries qualifiers worth reading before anyone drafts a migration plan. Artificial Analysis builds the AA-WER Streaming index from roughly eight hours of English audio drawn from AA-AgentTalk, VoxPopuli and Earnings22, and it does not touch the more than 70 languages Muse was trained on or the 25 Meta had verified at launch [7]. The margin over Cartesia's Ink-2 is 0.3 percentage points [2][3]. A third of a point on eight hours of English is a test result, not a product ranking, which is why AssemblyAI's separate benchmark seats a different model first [6]. The price is what survives that comparison, because it does not move when the leaderboard does.
The mechanism shows where Muse spends its advantage. It processes audio in 80-millisecond chunks and decides at each one whether to commit a word or wait for more context, with reinforcement learning multiplying a reward for lower error against a reward for shorter delay so that failing either measure disqualifies the choice [9][10]. The partial-transcript numbers make the cost visible: Muse returns 3.6% error at 0.13 seconds, while Cartesia Ink-2 with external endpoint detection returns 4% at 0.07 seconds [8]. Sixty milliseconds against four-tenths of a point [20] is a genuine trade, and the side you want depends on whether the transcript drives an agent's turn-taking or becomes a record a person reads later.
Google picked a different axis altogether. Gemini 3.5 Transcribe, launched August 26 and tested through Gboard's Rambler, strips verbal stumbles and converts spoken corrections into polished text [13]. Ryan Whitwam, the senior technology reporter who reviewed it, found the cleanup useful for short blocks of dictation while noting that the transcript was no longer exactly what he had said [14]. That cleanup helps drafting, but for anything that has to stand as the record of what someone actually said, a model that edits carries an exposure the per-hour rate does not capture.
Speaker labelling is the column where both sides should be modest. Meta puts its diarization error rate at 17.5% and says that led public diarization benchmarks at launch [5], a figure 5.6 times its own word error rate [18]. Muse handles more than 20 speakers and audio over an hour without a separate processing stage [11], so the pipeline gets simpler while attribution stays a review item, which matters for meeting minutes and call quality assurance that key on who spoke.
This quarter's decision is narrow: run your own audio through both and find out whether the difference lands as savings or as an evaluation bill. Next quarter's consequence deserves naming now. Muse ships through the Meta Model API and already powers system-wide dictation in Meta AI for Mac and Muse Code [12], so transcription is priced as an on-ramp rather than as a business. A buyer who moves on rate alone is accepting that the roadmap belongs to someone whose interest in speech-to-text is a means to something else.
Ranked by verification strength, evidence, and original report placement.
On the Artificial Analysis AA-WER Streaming benchmark displayed September 1, Muse produced a 3.1% word error rate on the final transcript 0.16 seconds after speech ended, narrowly ahead of Cartesia Ink-2 at 3.4%; the lead over Cartesia is 0.3 percentage points.
On the same benchmark, ElevenLabs Scribe v2 Realtime scored 3.6%, OpenAI's GPT Live Transcribe 3.9%, and Google's Gemini 3.5 Transcribe Live 4%.
Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, 2026, its first real-time model combining streaming transcription, speaker diarization and endpointing; one model transcribes words, separates speakers and detects when a person has stopped talking.
Muse costs $0.18 per hour, roughly 80% below Google Cloud Speech-to-Text's standard $0.96-per-hour rate, and is also cheaper than Cartesia, ElevenLabs and Deepgram Flux by smaller margins.
Artificial Analysis builds the AA-WER Streaming index from about eight hours of English audio drawn from AA-AgentTalk, VoxPopuli and Earnings22, and does not test Meta's performance across the more than 70 languages used in training, including the 25 languages Meta had extensively verified at launch.
On the first partial transcript after speech ended, Muse recorded 3.6% error at 0.13 seconds, Cartesia Ink-2 with external endpoint detection reached 4% at 0.07 seconds, and Cartesia's semantic endpoint setting returned 4.9% at 0.17 seconds.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Meta's coding agent has two prices: pay 18x more, or let it train on your repository1 distinct publisher
build
fal reports 35x MiniMax's own H3 endpoint after tuning the weights to its runtime1 distinct publisher
build
Harness choice moved token use 83-fold with the model held constant1 distinct publisher
build
Artificial Analysis moves eval onto your data, and turns model choice into procurement1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise numbers, one narrow test, one outlet
The figures are specific to two decimal places and the accuracy ones come from Artificial Analysis rather than from Meta, which is the strongest thing going for them. Against that: about eight hours of English audio stands behind a claim about a model trained on 70-plus languages, the diarization and language numbers are Meta describing Meta, and every one of these details reaches a reader through implicator.ai alone.
Shipping, all of it inside Meta
Muse is genuinely in production — Fn-key dictation on the Mac and Muse Code are real workloads, not demos — but both belong to Meta. There is no outside customer, no minutes-processed figure, and no third-party report of what the latency looks like outside a benchmark harness. Two days after launch that is availability, not pull.
The price is firmer than the crown
'First in streaming transcription' rests on three-tenths of a point over Cartesia across roughly eight hours of English, and AssemblyAI's board reportedly hands first place to someone else. Move to the earlier partial transcript and Cartesia is quicker. The 17.5% speaker-label error would not have led any press release. The overstatement sits in the launch framing, not in implicator.ai's write-up, which puts the thinness in its own subheading.
Vendor-chosen numbers on a vendor-friendly test
Two of the three strengths on offer — diarization leadership and the 25 verified languages — are Meta asserting things about Meta. The third is a leaderboard Meta can cite because it wins there by a rounding error, while the benchmark that does not flatter it appears in a single line with no numbers. Launching six days after Google, at a price that undercuts every specialist listed, is not a neutral act either.
Solid where it is arithmetic
The money holds up: $0.78 an hour, $780 a month, $9,360 a year on a thousand hours, all derived from published rates. The ranking does not hold up as well, because it depends on one test suite no second publisher here has examined and on Meta's word for the multilingual and speaker-labelling halves. Enough to justify a pilot; not enough to sign a multilingual contract on.