Leadership1 distinct publisher3 min readPublished
A 61 on Artificial Analysis's index no longer requires flagship pricing. That changes what agent workloads should cost this quarter. The same release also grew dearer than its own predecessor and slipped on two evaluations.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
Meta produced the 42% gap without touching its rate card [16]. Token prices held at the Muse Spark 1.2 levels: $1.25 per million input tokens, $4.25 per million output, and $0.15 per million for cached input [8]. Cached input therefore sits at 12% of the standard input price, an 88% discount on precisely the line item that moved, since Artificial Analysis attributed the entire per-task increase over version 1.2 to roughly 57% more input tokens against about 8% more output [3][5]. Gemini 3.8 Flash sat nearest in the same comparison at 59 points for $0.58, about 5% more per task for two fewer points, so the sub-dollar band holds more than one vendor [3][7]. What a buyer controls here is prompt shape and cache hit rate, not the price list.
The two available measurements of token use point in opposite directions. Meta's engineers measured about 20% fewer tool calls and about 25% fewer tokens for 1.3 than for 1.2 [12], while the index measured input use per task rising by more than half [5]. Both can hold on different task mixes, and implicator.ai notes that Meta's own comparison table runs the max variant on Meta's harness, which is not the same class of evidence as the independent index [10]. That leaves one figure measured the same way across all four vendors, the index's cost per task, and it is the only one a procurement conversation can stand on.
A skeptic would say a tie on a nine-evaluation composite is not a tie on any particular workload, and the record supports the objection [9]. Meta's table shows the split: Muse Spark led GPT-5.6 Sol on SWEAtlas CodeBase QnA and trailed it on DeepSearchQA and the Agentic IF Index [13]. What survives the objection is where the gains landed. Tau3-Bench Banking rose from 35% to 47% and Terminal-Bench 2.1 from 80% to 85% between the two releases [11], a twelve-point and a five-point move in multi-step tool use [5], which is the shape of the agent work most buyers are now repricing.
The top of the board sits behind a partner agreement. Max scored 62, one point above xhigh, and remains limited to Meta partners with no disclosed price [7], so the highest score a general customer can transact is 61 at 55 cents [8]. Mark Zuckerberg described the release on X as "frontier performance almost too cheap to meter" [14], and the same release costs 37.5% more per task than the version it replaced [2] while giving back four points on AA-LCR [6]. Both readings are accurate, and they describe the same operating condition: the price floats with model behaviour instead of holding to a posted rate. Repricing agent workloads on this evidence books a real saving this quarter and hands next quarter a budget line whose variance is set by someone else's token accounting.
Ranked by verification strength, evidence, and original report placement.
On September 2, Muse Spark 1.3 in xhigh reasoning mode scored 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high).
Cost per task on that index was $0.55 for Muse Spark 1.3, against $0.95 for GPT-5.6 Sol and $0.94 for Grok 4.6.
No model scoring at least 59 in the September 2 comparison had a lower cost per task than Muse Spark 1.3; Gemini 3.8 Flash was nearest, scoring 59 at $0.58.
Cost per task rose from $0.40 for Muse Spark 1.2 in August to $0.55 for version 1.3 on September 2; Artificial Analysis measured roughly 57% more input tokens per task and about 8% more output tokens, and attributed the higher cost to the heavier input use.
AA-LCR fell from 83% for Muse Spark 1.2 to 79% for both new reasoning modes, and AA-Omniscience accuracy dropped three percentage points for xhigh and one for max; Artificial Analysis attributed the decline to a higher abstention rate, which also lowered xhigh's hallucination rate.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Meta's coding agent has two prices: pay 18x more, or let it train on your repository1 distinct publisher
build
Muse Code's $5 tier buys the same 10 requests per five hours that Codex Plus starts at2 distinct publishers
product
Meta folds speaker labeling and endpointing into one streaming transcription model1 distinct publisher
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Third-party numbers, one messenger
The figures doing the work here — 61, $0.55, AA-LCR at 79% — come from Artificial Analysis rather than from Meta, which is the strongest thing this story has going for it. It is also the limit: one publisher has read those tables and nobody has checked the reading. implicator.ai keeps the tiers honest, putting Meta's harness table and its engineers' 20%-fewer-tool-calls count visibly below the independent index.
Shipped everywhere, used nowhere yet
What we can actually observe is a release date, two distribution channels and a benchmark run. No customer, no workload, no token volume, no migration off a rival — and the 62-point variant is still partners-only while safety testing finishes. Availability on day two is not adoption, and this reporting does not pretend otherwise.
"Too cheap to meter" versus a 37.5% price rise
Zuckerberg's phrase and the release it describes point in different directions: measured against its own predecessor, this model costs 37.5% more per task, and it gave back four points on long-context reasoning to get where it got. The overstatement lives in Meta's framing rather than in the coverage — implicator.ai runs the deflation in the same piece, headline included, which is why the gap is a lean and not a chasm.
Vendor harness, evaluator's scoreboard, newsletter pitch
Three interested parties shape what a reader sees. Meta benchmarks its partners-only variant on its own harness and has its chief executive supply the slogan. Artificial Analysis's standing rests on its index being treated as the scoreboard everyone tunes for. And implicator.ai sells a daily briefing on precisely this beat, with the signup dropped between the pricing section and the benchmark section. Only the first of the three is disclosed in the copy.
Precise, checkable, and checked once
Nothing here is vague — dated figures, named modes, stated attributions, arithmetic that survives a calculator. Nothing here is corroborated either. A second outlet reading the same September 2 index tables would move this number more than any amount of additional detail from Meta, and the softest link is the unstated question of whether the index's task mix resembles anyone's production traffic.