BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Same weights, 70 points apart: the ARC-AGI-3 table has stopped being procurement evidence
One model was verified at 30.16 in the official harness and reported at 100.00 in NVIDIA's. Microsoft's Agent Lightning now trains the harness into the weights.
The Engineer · Build desk

What happened
- ARC Prize verified Claude Opus 5 at 30.16 percent on the ARC-AGI-3 public set on July 24.
- NVIDIA reported 100.00 for the same weights on the same set on August 21, with only the surrounding code changed.
- OpenAI moved GPT-5.6 Sol from 13.3 to 38.3 percent by keeping reasoning and compacting instead of truncating, at a sixth of the output tokens.
- No perfect public-set score has been verified on the private set, and each author states as much.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A leaderboard row can no longer separate two vendors' weights when one set of weights covers most of the scale depending on who wrote the loop around it.
- exposure Fusing scaffolding into the checkpoint means a later model swap is a rebuild of the loop, not a configuration change, and the buyer owns that work.
- cost The points being bought sit in the layer the author expects to lose value with each frontier release, so the premium paid for them is a wasting asset.
- contradiction NVIDIA's own disclaimer that its comparison is not a controlled ablation removes the ground for reading the table as a ranking of models, which is how it is being read.
The scoring rule is what makes plumbing pay. ARC-AGI-3 scores each level as (human_baseline_actions / ai_actions) squared, with the ratio capped at 1.15 times the human baseline [7]. Squaring inflates a modest saving: AVO's headline result against VISTA is 12 percent fewer actions [8], and removing 12 percent of the actions lifts the ratio to about 1.136, which is roughly 1.29 once squared [21]. The cap permits 1.3225 [22]. One decision about how the loop spends its moves therefore collects around nine tenths of all the headroom the metric will ever give an agent that already matches a first-time human [23]. That is a target for whoever writes the loop, and it is invisible in a column headed by the model's name.
Put the harness papers side by side and the same parts recur under different labels, according to the dev.to tally. Memory: VISTA keeps a lossless record of every past observation, AVO carries forward prior implementations, evaluation results, compiler and profiler output and accumulated reasoning, and OpenAI's two flags are memory flags [9]. Supervision: AVO runs a monitor that watches for stagnation or repeated unproductive cycles and can redirect the main agent [10]. The official harness went the other way, discarding private reasoning after every game action and truncating history on a rolling window [11], because ARC built it to be generic and to make model shortcomings more visible [12]. The bottom row measures a model under a harness designed to expose it, the top rows measure models under harnesses designed to cover for them, and nobody has isolated which part of the gap is which [19].
The vendors are more careful than their readers. NVIDIA states that the AVO-versus-VISTA comparison should not be interpreted as a controlled ablation, and that the results are not a direct measurement of AVO's contribution [14]. VISTA notes its models were released after the public games, that overlap cannot be excluded, and that the private set remains the real test of generalisation [15]. Schema makes no frozen-harness or held-out-performance claim [16]. MIT reproduced the effect on August 5 and a group led by Impossible Research reached 98.98 on July 15 [17]; Google added a framework that gives the environment a harness of its own [26].
Once reinforcement learning runs with the deploy-time harness owning the loop, as Microsoft's Agent Lightning v1.0 does, the harness stops being a wrapper you can strip off for comparison and starts being part of the weights [1]. The dev.to author reads Agent Lightning's own reward-hacking section as the checklist his July post warned about [25]. A buyer evaluating two checkpoints trained that way is not comparing models at all, and there is no configuration change that will make them comparable.
Which leaves the depreciation question the writer raised himself. His July split put compensatory layers, the ones patching what a model cannot yet do, in the pile he expected to lose value with every model release, and all three components above sit in that pile [24]. They are currently worth 25 to 70 points against the newest frontier models on a benchmark built to resist static tricks [6]. Either the prediction was early or it was wrong. Either way, his procurement test is the usable output: a score without a harness version, a memory state and an action budget attached is a self-reported claim [2]. On that reading, "verified at 30.16" [3] certifies a run, not a model.
What to watch
- Private-set verification of any near-ceiling public-set score, which would show how much of the gap is memory and supervision and how much is familiarity with the 25 public games.
- Whether ARC Prize starts publishing harness version, memory state and action budget alongside each verified row.
- Whether an Agent Lightning-trained checkpoint keeps its score when moved back into the official generic harness.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence54
- Adoption58
- Hype gap+28
- Incentives66
- Confidence48
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Microsoft published Agent Lightning v1.0 on August 18; it runs reinforcement learning with the deploy-time harness owning the loop, so the harness becomes part of the weights.
- [2]
The author's test: a benchmark number without a harness version, memory state and action budget attached is a self-reported claim with an unmarked type.
- [3]
On July 24, ARC Prize verified Claude Opus 5 at 30.16% on the ARC-AGI-3 public set.
- [4]
On August 21, NVIDIA reported the same model at 100.00 on the same ARC-AGI-3 public set; the weights did not change, the code around them did.
- [5]
Retaining reasoning and enabling compaction took GPT-5.6 Sol from 13.3% to 38.3% and cut output tokens by 6x; OpenAI tripled the score on July 29 by flipping two API settings.
- [6]
On ARC-AGI-3's public set the spread between model in the official harness and model in the best harness is 25 to 70 points, on a benchmark designed to resist exactly this.
- [7]
ARC-AGI-3 scores agents with RHAE: per level, score = (human_baseline_actions / ai_actions)^2, with the ratio capped at 1.15x the human baseline. Game scores are level-weighted averages, the last level must be finished for full credit, and the overall number is the mean over games.
- [8]
AVO's headline result against VISTA is 12% fewer actions, which the author calls a harness optimisation target rather than a model property.
- [9]
VISTA keeps a lossless visual memory of every past observation; AVO carries forward prior implementations, evaluation results, compiler and profiler outputs and accumulated reasoning; OpenAI's two settings are memory settings (keep the reasoning, compact instead of truncate).
- [10]
AVO runs a monitor that watches the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent.
- [11]
The official ARC harness discarded all private reasoning after each game action and used a rolling truncation window, so older actions vanished as history grew.
- [12]
OpenAI's write-up quotes ARC's intent: an intentionally generic harness, without tools or special features, built to make model shortcomings more visible.
- [13]
None of the 100.00 scores are verified on the private set, and every author says so; every 100 on the table is a public-set number on games the models may have seen in training.
- [14]
NVIDIA states the AVO-versus-VISTA comparison should not be interpreted as a controlled ablation, and that the results should not be interpreted as a direct measurement of the performance contribution of AVO.
- [15]
VISTA states its models were released after the public ARC-AGI-3 games, that overlap cannot be excluded, and that the private set remains the real test of generalization.
- [17]
MIT did the same thing on August 5, and a group led by Impossible Research reached 98.98 on July 15.
- [18]
A 100.00 on ARC-AGI-3 means the agent finished every level at least as efficiently as a first-time human, across 25 games.
- [19]
The leaderboard measures model plus a harness built to expose the model, the 100s measure model plus a harness built to cover for the model, neither isolates the model, and nobody has established which part of the 70 points is which.
- [20]
The difference between the verified official-harness score and NVIDIA's reported score for the same weights is 69.84 points.
- [21]
Cutting actions by 12% raises the human/AI action ratio to about 1.136, which squared gives a per-level score of about 1.29, or 29% above parity with the human baseline.
- [22]
The 1.15x ratio cap limits any per-level RHAE score to 1.3225, that is 32.25% above parity.
- [23]
A 12% action saving captures about 90% of the headroom the RHAE cap allows above parity.
- [24]
The author's July taxonomy split harness work into compensatory layers that patch what the model cannot do yet and protective layers that constrain what it must not do, predicted the compensatory pile depreciates with every model release, and places memory, supervision and action budget in that pile.
- [25]
The author reads Agent Lightning's reward-hacking section as the checklist his July post warned about.
- [26]
On August 20, Google published a framework that gives the environment a harness of its own.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toThe Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- ARC-AGI-3 EvaluationFollow
- Benchmark ContaminationFollow
- Agent Harness EngineeringFollow
- Agent Benchmark IntegrityFollow
- Harness-in-the-Loop Reinforcement LearningFollow
- Eval Provenance and ProcurementFollow
Entities
- ARC-AGI-3Follow
- Relative Human Action Efficiency (RHAE)Follow
- ARC PrizeFollow
- Claude Opus 5Follow
- GPT-5.6 SolFollow
- NvidiaFollow
- Agentic Variation Operators (AVO)Follow
- VISTAFollow
- OpenAIFollow
- MicrosoftFollow
- Agent Lightning v1.0Follow
- GoogleFollow
- MITFollow
- Impossible ResearchFollow
- SchemaFollow
- SWE-bench VerifiedFollow
- Qwen3.5-9BFollow
- SWE-smithFollow
- mini-SWE-agentFollow