Science1 distinct publisher3 min readPublished
ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness and 99.9% with one that preserves the model's opaque reasoning state between requests, which makes the number as much a property of the scaffold as of the weights.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The two conditions differ in what the model is permitted to remember. ARC Prize's standard harness gives Astra continuity only through notes it writes to itself and chooses to carry forward, so anything it does not write down is absent on the next request [2]. The provider adapter harness preserves the model's opaque reasoning state between requests and compacts the conversation as it lengthens, letting prior work be reused rather than rebuilt [4]. The distance between the two results is 37.2 points, 99.9 minus 62.7 [15].
Suppose most of that spread belongs to the harness, which is the reading the write-up invites. Then 37.2 points is an estimate of how much a frontier model loses when it has to compress its own working state into prose for its future self. The replays show what survived that compression: object positions, coordinates, rules, unfinished plans, and a shorthand Astra generated per environment, with lines such as "extend8 to3; retract10 to2; shorten8 to1" standing in for a multi-step plan [14]. Dense self-summarization, and on this evidence still lossy against whatever the provider keeps internally [4].
The size of the spread is what complicates the board. Astra's standard-harness 62.7% sits 32.5 points above the 30.2% ARC Prize reports for Claude Opus 5 [16], which itself sits well above GPT-5.6 Sol at 7.8% [7]. So the gap between one model's two entries is wider than the gap between the top two systems [17].
The human costs in the same write-up are a lesson in denominators. Participants earned $115 for a 90-minute session plus $5 per completed game and attempted about nine games, roughly $12.78 per attempted game before bonuses [12]. ARC Prize also prices the brain's electricity alone at about 0.067 cents per game attempted, noting that most of the fee buys a participant's time and willingness rather than their energy [13]. Those two ways of costing identical human performance differ by a factor near 19,000 [21]. The agent side has no comparable per-game figure at all: both dollar totals cover a whole benchmark run, and no task count is published beside them [19].
The action-efficiency result is narrower than it first sounds. Astra used fewer actions than the median tested human on 96% of levels [8], which counts moves spent, not levels finished. And the thing this doesn't tell you is how far these environments carry into the agent loop you actually run: they are turn-based, withhold instructions, and require the agent to infer its own goals [10], which is the setting where uninterrupted private reasoning should pay most. With the leader now a tenth of a point from the ceiling [23] and humans solving every environment [9], ARC-AGI-3 has little room left to distinguish a good agent from a good harness.
Ranked by verification strength, evidence, and original report placement.
ARC Prize labels the 62.7% Standard harness result as Astra (max) and the 99.9% Provider Adapter result as Astra (high), and the material does not define what max and high denote.
OpenAI's GPT-6 Astra, labelled Astra (max), scores 62.7% on ARC-AGI-3 Semi-Private for $26K using ARC Prize's Standard harness.
ARC Prize's Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment.
With a new Provider Adapter harness, GPT-6 Astra, labelled Astra (high), scores 99.9% on ARC-AGI-3 Semi-Private for $19K.
The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
build
Same weights, 70 points apart: the ARC-AGI-3 table has stopped being procurement evidence1 distinct publisher
invest
Compute scarcity meters the model OpenAI says can fill out forms at superhuman speed1 distinct publisher
build
Nvidia says the harness, not the model, took Claude Opus 5 from 30.2% to 100% on ARC-AGI-34 distinct publishers
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise numbers, single scorer
Both scores, both dollar totals, the 96% action-efficiency figure and the two rival results all come from one post by the organisation that owns the benchmark and ran the test. What lifts this above the usual announcement is that ARC Prize describes the two harnesses and publishes the weaker number alongside the stronger one. What caps it is that the set is semi-private, so nobody outside can rerun it, and the harness that produced 99.9% is documented in a paragraph of prose.
Scoreboard only
Nothing here says anyone is running this model in anger. The whole record is two evaluation runs and two comparison figures, all posted by the benchmark's keeper: no deployments, no usage disclosures, no pricing beyond what the runs cost ARC Prize. Turning a leaderboard row into uptake would be our invention, not this reporting's.
The scaffold is doing some of the winning
A 99.9% on a benchmark humans solve completely reads as finished, and ARC Prize calls both runs state-of-the-art. Yet the same weights score 62.7% when the harness only lets the model keep notes it wrote itself, so 37 of those points are attributable to machinery that carries hidden reasoning state between requests. ARC Prize does not bury this - it is in the same paragraph - but it also never defines max and high, never gives the task count that would make $19K mean something per game, and never says how the models it compares against were run.
Both parties gain from the same headline
ARC Prize designed the benchmark, ran the evaluation, wrote the adapter behind the top score, and gains standing every time its residual-gap framing anchors a frontier claim; OpenAI gets a 99.9%. The human cost comparison shows the pull plainly - the same paid game attempt is priced at $12.78 and at 0.067 cents, some nineteen thousand times apart, and the figure the post advances as the fairer proxy is the one that makes the machine look dear or cheap by choice of denominator.
Sure what was said, unable to check it
The arithmetic is solid and the claims are unambiguous, so we are confident about what was asserted and how the numbers relate. We are much less confident that the assertions hold, having exactly one first-party account and no way to replicate a semi-private evaluation. Our read is that the harness disclosure here is unusually candid and the framing built on top of it is not.