Skip to content

Product1 publisherNot yet confirmed elsewhere3 min readPublished

Nvidia moved one model from 30% to 100% without changing the model

The harness did the work on ARC-AGI-3, which puts agent reliability on your engineering backlog rather than in your vendor contract.

The Product Desk

How we use AISend a correction

Photograph accompanying Nvidia moved one model from 30% to 100% without changing the model
Photo: anthropic.com

What happened

  • Nvidia researchers put Claude Opus 5 behind a custom harness with memory handling and a supervisor component and scored 100% on ARC-AGI-3.
  • The same model without that harness scored 30%, which was still the best result of any model tested bare.
  • OpenAI, whose models scored under 10%, tripled its own results last month by changing two harness settings and did not reach 100%.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Reliability work moves off the vendor evaluation spreadsheet and onto the engineering backlog, because the larger share of the measured gain came from code the customer would have to write.
  • cost If a harness can double spend on an identical model, per Databricks, then internal model cost comparisons run across teams with different harnesses are measuring the wrong thing and someone is...
  • constraint Teams standardised on a single-layer coding agent have no supervisor tier to add without building one, so the ceiling in the paper is not reachable by upgrading a subscription.
  • contradiction Both labs report roughly threefold gains from harness changes, but only Nvidia reached the ceiling, which leaves the credit split between the supervisor design and Opus 5 unsettled for anyone...

Seventy of those hundred points came from software the researchers wrote, not from the model they licensed. Claude Opus 5 produced 30 on its own, and that was the best any bare model managed in the test [3]. So model selection could only ever have moved the result across a 30-point spread there, while the scaffolding moved it by 70 [15]. Anyone comparing frontier models on an agentic workload is arguing over the smaller variable.

The specific piece Nvidia credits is a second agent watching the first. Adel El Hallack, the vice president of product in Nvidia's AI unit, described it as nudging the worker agent when it drifts, heads into a dead end, or starts re-walking ground it already covered [6]. The idea is not new [13], but most teams today run a single harness layer, whether that is Claude Code, Codex or Hermes [7]. The gap between one layer and two is not something a vendor is currently selling you.

It is also a cost variable, not only a quality one. Databricks found in July that the harness, more than the model, drives spend, and Ali Ghodsi's version of it is blunt: the same model on the wrong harness can double your bill [11]. Which means the model price comparison your finance team is running across two teams with two different harnesses is not a comparison at all.

Two caveats sit inside the result. ARC-AGI-3 is a set of 2D games with no instructions, where a 100% score means playing as well as a human [4]. A supervisor can tell when the worker is stuck because there is a score that stops moving. Most long-horizon business work has no such signal, and Microsoft's April test of 19 LLMs on document editing, where every model including the frontier ones produced error-filled output [10], is the reminder of what happens without one. The second caveat is attribution. OpenAI's own ARC-AGI-3 work tripled its scores by changing two harness settings and still fell short of 100 [9]; Nvidia's jump from 30 to 100 is a 3.3x move, arithmetically close to OpenAI's 3x [14]. The direction agrees across both labs. Whether the last stretch belongs to the supervisor or to Opus 5 does not separate out from what has been published.

Note too who benefits from the finding. Nvidia's harness, Agentic Variation Operators, is not a product, but the company does ship harness components under the Nemo brand, some commercial [8]. A result locating the value in scaffolding suits a supplier of scaffolding parts. It is corroborated anyway. And the failure modes that come with agents running loose, including deleted files, deleted databases, and agents that resort to collusion or hacking to finish a job [12], land on whoever owns the loop. On Nvidia's own framing of what an agent is [5], that is your team.

What to watch

  • Whether the supervisor loop from AVO appears as a usable Nemo component rather than staying a published result.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories