Build6 distinct publishers3 min readPublished
Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Artificial Analysis weights GDPval-AA v2 at 20 percent of its Intelligence Index, Terminal-Bench 2.1 at 16 percent and tau3-Bench Banking at 14 percent, and Meta's largest gains landed in exactly those three tests [9]. Half the composite comes from those three tests [3]. The 61 is largely a statement about driving tools in a loop. On tau3-Banking, version 1.2 sat at 35 percent; the xhigh tier you can call today reaches 47, tying Claude Fable 5.1 max and GLM-5.3-Flash, and the preview max tier hits 52, which Artificial Analysis calls the only outright lead the model holds [10]. Terminal-Bench 2.1 climbs from 80 to 85 on xhigh [11]. Claude Fable 5.1 is still at 91.4 in its max tier, so the generally available Spark trails the leader by 6.4 points on the terminal coding test [4]. On the highest-weighted test, Meta moves 1,615 to 1,709 at xhigh against Claude Fable's 1,853, on a scale calibrated to human expert performance at 1,000 across 220 professional tasks [12] [5].
The budget arithmetic is cleaner. Input and output prices did not move: $1.25 and $4.25 per million tokens [8]. The cost of one index task did, from $0.40 on 1.2 to $0.55 [8]. Same per-token rate, 37.5 percent more money per task, which means roughly 37.5 percent more tokens burned per task [1]. Reasoning models buy scores with output. Rivals at the same index level run $0.94 to $1.23, so Spark lands between 1.7 and 2.2 times cheaper per task [2], and Artificial Analysis calls it the most cost-efficient model at its measured intelligence, on its Pareto line [19]. That is good engineering and worth saying plainly. Meta buys the max variant's single extra index point with 62 percent more reasoning tokens, which is one way to spend a point [13] [7].
What would have to be true for Meta's own table to describe your workload. The DeepSWE 75.4 percent and the 98.1 on the 512K-1M MRCR split were assembled by Meta, with rival figures drawn from a mix of its own evaluations, leaderboards and vendor-reported results, and its testing of third-party models described as best-effort [4] [5]. The headline comparisons ran at max, and xhigh is the highest level developers can call [6]. Your tasks would need to look like tool sequences against a verifier, and you would need access Meta has not opened. Meanwhile AA-LCR fell from 83 to 79 percent, and factual accuracy on AA-Omniscience slipped by up to three points because the model declines more often when unsure [15]. Long-context retrieval work gets a slightly worse model with a better leaderboard position.
The New Stack headlined its coverage with 1.3 edging out Gemini [18]. On GPQA Diamond, the same Artificial Analysis run has Gemini 3.8 Flash high at 95.3 against Muse Spark's 94 [14]. Both readings are available depending on which test you weight, which is the argument against standardizing a toolchain on any one vendor: 1.2 shipped in August, 1.3 arrived less than a month later, and Zuckerberg has already teased a larger model codenamed Watermelon [1] [2] [17]. The durable move is a vendor-neutral eval harness and tool schemas, with the model bill treated as a monthly line item rather than a fixed bet. The promised open-weight release would loosen the coupling further, but Meta has not published the licence, and its previous open-weight models carried varying restrictions on use [16].
Ranked by verification strength, evidence, and original report placement.
Muse Spark 1.3 pricing is unchanged at $1.25 and $4.25 per million input and output tokens, and one Intelligence Index task costs $0.55; no model scoring 59 or higher is cheaper, rivals at the same index level run between $0.94 and $1.23, and version 1.2 ran $0.40 per task.
Meta released Muse Spark 1.3 through Muse Code and the Meta Model API; it is the company's fourth model in five months, with the series launching in April, 1.1 following in July and 1.2 in August.
Muse Spark 1.3 arrives less than a month after the release of Muse Spark 1.2.
Meta ran Spark 1.3 at its new "max" reasoning level for the headline comparisons, while the highest reasoning level generally available to developers is "xhigh"; max remains in limited preview while Meta completes additional safety testing.
Artificial Analysis scores the publicly available Muse Spark 1.3 xhigh at 61 on its Intelligence Index, four points ahead of Spark 1.2 and level with GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high; the limited-preview max version scored 62, behind only Claude Fable 5.1 and Claude Opus 5 in its comparison at launch.
Artificial Analysis describes Spark 1.3 xhigh as the "most cost-efficient model" at its level of measured intelligence, placing it on the company's Pareto line.
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
Meta's Muse Spark 1.3 matches three flagship models at 55 cents a task1 distinct publisher
product
Meta keeps Muse Spark 1.3 pricing flat while claiming coding edge over GPT-5.63 distinct publishers
invest
Meta's coding agent has two prices: pay 18x more, or let it train on your repository1 distinct publisher
build
Gemini 3.8 Flash's introductory price doubles on December 31, 20267 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
blog.vercel.com
1 article · September 1, 2026
infoworld.com
1 article · September 3, 2026
latent.space
1 article · September 2, 2026
runtimewire.com
1 article · September 2, 2026
the-decoder.com
2 articles · September 3, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Independent numbers, one independent scorer
Unusually for a launch story, the load of the argument sits outside the vendor: Artificial Analysis tested both tiers, published per-test detail, and reported results Meta would not have chosen to lead with — AA-LCR down four points, factual accuracy slipping, a single outright lead. What keeps this short of strong is concentration. Every score in our coverage traces back to that one benchmarking firm or to Meta's own best-effort table, and the tier in the headline comparison is not one a developer can call.
Live on real rails, no usage behind it
The distribution is genuine rather than announced: Muse Code and the Meta API on day one, Vercel's Gateway listing it the same day, selectable inside Claude Code, Codex and Cursor. What is missing is anyone using it. The single quantity in the entire story is Wang's own "meaningful double digit" percentage of coders opting into the training-discount tier — a figure with no denominator and no external check — and no named enterprise deployment appears anywhere.
Frontier framing, preview-tier proof
"Frontier performance almost too cheap to meter" is measured, in the same week, as one outright benchmark lead — on a simulated banking scenario, at a reasoning tier in limited preview. The shipping tier ties rather than leads there, trails Claude Fable 5.1 on both heavyweight tests, gave back four points on long-context retrieval, and costs 37.5 percent more per finished task than the version it replaces while Meta advertises using fewer tokens. Half the index weight happens to sit in the three tests where the gains landed. The price story is the part that survives contact with the numbers; the frontier story is running ahead of them.
Everyone in frame is selling something
The comparison table is written by the company being compared, from runs it calls best-effort, at a tier it has not released. The scoreboard everyone quotes belongs to a commercial benchmarking firm whose weighting choices decide the headline. The changelog telling builders how to switch models comes from the gateway that bills for the tokens. The cheerleading quote comes from a cloud provider's evangelist, the caution from a rival AI product's blog, and the discount tier exists because Meta wants the training data. Only The Decoder's audit and InfoWorld's analyst panel are positioned to lose nothing either way.
Solid where it counts, thin at the edges
The pricing, the index scores and the per-test detail are consistent wherever two publishers touch the same fact, and the caveat about the preview tier is repeated independently by three of them. Confidence drops on the material only one publisher carries: the regressions and the index weightings come solely from The Decoder, the Gemini displacement and the Watermelon tease solely from The New Stack, and that tease rests on an emoji plus a July report. Worth noting too that our six publishers include two whose stories appear twice, so the apparent breadth is narrower than the count suggests.
thenewstack.io
2 articles · September 3, 2026