Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Same weights, 70 points apart: the ARC-AGI-3 table has stopped being procurement evidence

One model was verified at 30.16 in the official harness and reported at 100.00 in NVIDIA's. Microsoft's Agent Lightning now trains the harness into the weights.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Same weights, 70 points apart: the ARC-AGI-3 table has stopped being procurement evidence
Generated illustration

What happened

  • ARC Prize verified Claude Opus 5 at 30.16 percent on the ARC-AGI-3 public set on July 24.
  • NVIDIA reported 100.00 for the same weights on the same set on August 21, with only the surrounding code changed.
  • OpenAI moved GPT-5.6 Sol from 13.3 to 38.3 percent by keeping reasoning and compacting instead of truncating, at a sixth of the output tokens.
  • No perfect public-set score has been verified on the private set, and each author states as much.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A leaderboard row can no longer separate two vendors' weights when one set of weights covers most of the scale depending on who wrote the loop around it.
  • exposure Fusing scaffolding into the checkpoint means a later model swap is a rebuild of the loop, not a configuration change, and the buyer owns that work.
  • cost The points being bought sit in the layer the author expects to lose value with each frontier release, so the premium paid for them is a wasting asset.
  • contradiction NVIDIA's own disclaimer that its comparison is not a controlled ablation removes the ground for reading the table as a ranking of models, which is how it is being read.

The scoring rule is what makes plumbing pay. ARC-AGI-3 scores each level as (human_baseline_actions / ai_actions) squared, with the ratio capped at 1.15 times the human baseline [7]. Squaring inflates a modest saving: AVO's headline result against VISTA is 12 percent fewer actions [8], and removing 12 percent of the actions lifts the ratio to about 1.136, which is roughly 1.29 once squared [21]. The cap permits 1.3225 [22]. One decision about how the loop spends its moves therefore collects around nine tenths of all the headroom the metric will ever give an agent that already matches a first-time human [23]. That is a target for whoever writes the loop, and it is invisible in a column headed by the model's name.

Put the harness papers side by side and the same parts recur under different labels, according to the dev.to tally. Memory: VISTA keeps a lossless record of every past observation, AVO carries forward prior implementations, evaluation results, compiler and profiler output and accumulated reasoning, and OpenAI's two flags are memory flags [9]. Supervision: AVO runs a monitor that watches for stagnation or repeated unproductive cycles and can redirect the main agent [10]. The official harness went the other way, discarding private reasoning after every game action and truncating history on a rolling window [11], because ARC built it to be generic and to make model shortcomings more visible [12]. The bottom row measures a model under a harness designed to expose it, the top rows measure models under harnesses designed to cover for them, and nobody has isolated which part of the gap is which [19].

The vendors are more careful than their readers. NVIDIA states that the AVO-versus-VISTA comparison should not be interpreted as a controlled ablation, and that the results are not a direct measurement of AVO's contribution [14]. VISTA notes its models were released after the public games, that overlap cannot be excluded, and that the private set remains the real test of generalisation [15]. Schema makes no frozen-harness or held-out-performance claim [16]. MIT reproduced the effect on August 5 and a group led by Impossible Research reached 98.98 on July 15 [17]; Google added a framework that gives the environment a harness of its own [26].

Once reinforcement learning runs with the deploy-time harness owning the loop, as Microsoft's Agent Lightning v1.0 does, the harness stops being a wrapper you can strip off for comparison and starts being part of the weights [1]. The dev.to author reads Agent Lightning's own reward-hacking section as the checklist his July post warned about [25]. A buyer evaluating two checkpoints trained that way is not comparing models at all, and there is no configuration change that will make them comparable.

Which leaves the depreciation question the writer raised himself. His July split put compensatory layers, the ones patching what a model cannot yet do, in the pile he expected to lose value with every model release, and all three components above sit in that pile [24]. They are currently worth 25 to 70 points against the newest frontier models on a benchmark built to resist static tricks [6]. Either the prediction was early or it was wrong. Either way, his procurement test is the usable output: a score without a harness version, a memory state and an action budget attached is a self-reported claim [2]. On that reading, "verified at 30.16" [3] certifies a run, not a model.

What to watch

  • Private-set verification of any near-ceiling public-set score, which would show how much of the gap is memory and supervision and how much is familiarity with the 25 public games.
  • Whether ARC Prize starts publishing harness version, memory state and action budget alongside each verified row.
  • Whether an Agent Lightning-trained checkpoint keeps its score when moved back into the official generic harness.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence54
Adoption58
Hype gap+28
Incentives66
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Microsoft published Agent Lightning v1.0 on August 18; it runs reinforcement learning with the deploy-time harness owning the loop, so the harness becomes part of the weights.

  2. [2]

    The author's test: a benchmark number without a harness version, memory state and action budget attached is a self-reported claim with an unmarked type.

  3. [3]

    On July 24, ARC Prize verified Claude Opus 5 at 30.16% on the ARC-AGI-3 public set.

    ReportedSupportedSource: dev.to post aggregating ARC-AGI-3 resultsView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 24, 2026

    The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories