Product1 publisher3 min readPublished
ARC Prize ran GPT-6 Astra under two scaffolds and published both, 62.7% through its own minimal interface and 99.9% through OpenAI's Provider Adapter, which cost less. An eval that leaves the harness loose is scoring plumbing.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
With reasoning effort set to none inside OpenAI's Provider Adapter, Astra scores 96.7%, which is 34 points above the same weights running at maximum reasoning inside ARC Prize's standard harness [6]. The setting vendors sell as the capability lever lost to a configuration choice about state management.
That choice is small and legible. ARC Prize's standard harness gives every model the same minimal interface and lets the model decide which notes to carry forward, while OpenAI's adapter preserves the model's opaque reasoning state between requests and compacts longer conversations [5]. Whether the model keeps its own scratchpad, and who trims context when it fills, is the entire delta between the two numbers.
The cheaper run was also the better one. The standard harness cost $26,098 for 62.7% [2]; the adapter cost $18,817 for 99.9% [3]. That is $416 per point against $188 per point [2], and the losing run cost 38.7% more in absolute dollars [3]. On the 167 game-reasoning pairs both harnesses solved, the adapter used 49% fewer tokens and ran roughly 3.66 times faster [7]. Scaffolding shows up on the capability line and on the invoice at the same time.
Two limits apply to how far that finding travels. This is one benchmark family and one model, so nothing here shows the same harness dominance on retrieval or support triage. And the pinned figure is not perfectly pinned: Francois Chollet posted 66% where ARC Prize's table says 62.7%, a 3.3-point spread depending on which artifact you read, and the blog is the primary record [12][4]. ARC Prize says it will publish both harness results side by side from now on [11].
The rest of the launch numbers moved as well, and each metric was altered on its own terms. Fortune's Emily Forlini compared archived snapshots of OpenAI's launch post and found five metrics altered after publication [13]. Astra's hallucination rate read 4.2%, then 2% by 5.20pm, then 4.2% again [14]. Anthropic's Fable 5.1 fell from 87.8% to 78% on FrontierMath before settling at 83% [15], and Sol's ExploitBench doubled from 5.5% to 11.5% at a reasoning level Sol does not offer commercially, which OpenAI told Fortune it is investigating reverting [16]. The embargo draft put ARC-AGI-3 at 98.6% against 99.99% in the live post [17]. Not every edit favoured Astra, since two Anthropic HealthBench Professional scores went up [23]. OpenAI's account to Fortune is noise of a few percentage points from checkpoint, scaffold and evaluation run [18], and Snorkel AI's Vincent Sunn Chen told Fortune that scores routinely shift in the final hours before a launch [21]. Anka Reuel and Mike Hardy, the Stanford researchers who call this benchmaxxing, went to the system card for method and found barely any details about the internal hallucination evaluation, including no count of test items [19][20].
For an eval spec, the forcing function is four fields filled for both sides of every comparison: which harness ran it, whether reasoning state persists between requests, who compacts context and at what threshold, and what reasoning effort was set to [4]. Then draw the grid, same weights against different weights on one axis, same harness against different harness on the other. Only the same-harness cells say anything about a model. The different-harness cells price your platform team's work, which is worth knowing and is not a model claim. The comparison that travelled, 99.9% against Sol's 7.8%, sits in the fourth cell with both variables loose [8], and the publisher that later documented the gap had carried that pairing without the caveat in its own launch coverage [9].
Ranked by verification strength, evidence, and original report placement.
ARC Prize published both a 62.7% and a 99.9% score for GPT-6 Astra on the day the model launched, along with a full table of every reasoning level it ran.
Under ARC Prize's standard harness at maximum reasoning, GPT-6 Astra scored 62.7% and the run cost $26,098.
Under OpenAI's Provider Adapter at high reasoning, GPT-6 Astra scored 99.9% and the run cost $18,817.
A harness is the software around a model: it sets the tools the model can reach, what it remembers between requests, and how its context gets managed.
ARC Prize's standard harness gives every model the same minimal interface and lets the model decide which notes to carry forward; OpenAI's Provider Adapter preserves the model's opaque reasoning state between requests and compacts longer conversations.
With reasoning effort set to none inside OpenAI's adapter, Astra still scored 96.7%, beating the same model at maximum reasoning inside the standard harness by 34 points.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Firm on the table, borrowed on the edits
Two grades of sourcing sit side by side. The harness gap comes from ARC Prize's own published table, with reasoning levels, dollar costs and token counts quoted and named statements from Knoop and Chollet attached, and it stands on its face. The revision trail is thinner in our hands: every altered metric reaches us through Fortune's archive comparison as relayed by The Next Web, and the 3.3-point disagreement between Chollet's post and the table is reported without resolution.
Shipped, priced and independently scored
Astra is live and has a public price, and two outside parties have already run it: ARC Prize across both scaffolds and Artificial Analysis on its own coding and intelligence indices, which it rebuilt a day later with private test sets. What this reporting never shows is the model doing work at a customer, so the uptake on record is measurement rather than deployment.
The headline needs the vendor's adapter
OpenAI's AGI declaration and the 99.9%-against-7.8% pairing that spread both depend on the run inside OpenAI's own adapter. The benchmark's like-for-like figure is 62.7% against 7.8%, still a large jump but a different story, and ARC Prize put in writing that it is not claiming AGI. Artificial Analysis pushes the same direction, scoring Astra level with the model it replaces on general intelligence at two and a half times Sol's price.
Every party in frame has a stake in the number
OpenAI wrote the launch post, revised five figures in it, pulled it and would not say why, and supplied the adapter that produced the higher score. ARC Prize is grading a vendor submission against two scaffolds of its own choosing while running a prize programme. Snorkel AI and Artificial Analysis both sell evaluation work, and The Next Web is marking its own earlier homework in public, which is disclosure and self-interest at once.
One publisher, one primary record
Where the numbers come straight from the benchmark's published record, they can be relied on. The revision timeline, the Stanford reading of the system card and OpenAI's noise explanation all arrive through a single publication's account of Fortune's reporting, The New Stack's independent side-by-side is mentioned rather than shown, and a figure inside ARC Prize's own house does not match itself.
invest
OpenAI's post-launch edits doubled Astra's math lead over Anthropic's Fable1 publisher
invest
Astra's 99.9% holds up only on the harness OpenAI ran itself1 publisher
invest
Sanders and Casar attach a 20-year prison term to building superintelligence1 publisher
build
Epoch's first-place ranking for GPT-6 Astra rests on a single coding score2 publishers
Publishers with included, body-backed reporting in this cluster.
1 article · September 6, 2026