Product1 distinct publisher3 min readUpdated
The harness did the work on ARC-AGI-3, which puts agent reliability on your engineering backlog rather than in your vendor contract.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Seventy of those hundred points came from software the researchers wrote, not from the model they licensed. Claude Opus 5 produced 30 on its own, and that was the best any bare model managed in the test [3]. So model selection could only ever have moved the result across a 30-point spread there, while the scaffolding moved it by 70 [13]. Anyone comparing frontier models on an agentic workload is arguing over the smaller variable.
The specific piece Nvidia credits is a second agent watching the first. Adel El Hallack, the vice president of product in Nvidia's AI unit, described it as nudging the worker agent when it drifts, heads into a dead end, or starts re-walking ground it already covered [6]. The idea is not new [15], but most teams today run a single harness layer, whether that is Claude Code, Codex or Hermes [7]. The gap between one layer and two is not something a vendor is currently selling you.
It is also a cost variable, not only a quality one. Databricks found in July that the harness, more than the model, drives spend, and Ali Ghodsi's version of it is blunt: the same model on the wrong harness can double your bill [11]. Which means the model price comparison your finance team is running across two teams with two different harnesses is not a comparison at all.
Two caveats sit inside the result. ARC-AGI-3 is a set of 2D games with no instructions, where a 100% score means playing as well as a human [4]. A supervisor can tell when the worker is stuck because there is a score that stops moving. Most long-horizon business work has no such signal, and Microsoft's April test of 19 LLMs on document editing, where every model including the frontier ones produced error-filled output [10], is the reminder of what happens without one. The second caveat is attribution. OpenAI's own ARC-AGI-3 work tripled its scores by changing two harness settings and still fell short of 100 [9]; Nvidia's jump from 30 to 100 is a 3.3x move, arithmetically close to OpenAI's 3x [14]. The direction agrees across both labs. Whether the last stretch belongs to the supervisor or to Opus 5 does not separate out from what has been published.
Note too who benefits from the finding. Nvidia's harness, Agentic Variation Operators, is not a product, but the company does ship harness components under the Nemo brand, some commercial [8]. A result locating the value in scaffolding suits a supplier of scaffolding parts. It is corroborated anyway. And the failure modes that come with agents running loose, including deleted files, deleted databases, and agents that resort to collusion or hacking to finish a job [12], land on whoever owns the loop. On Nvidia's own framing of what an agent is [5], that is your team.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Nvidia published research on Friday suggesting the harness, more than the underlying model, is far more important when asking an AI to do long-horizon tasks.
Using a custom harness tweaked to handle memory well and including a supervisor component, Nvidia researchers got Claude Opus 5 to a 100% score on the interactive reasoning benchmark ARC-AGI-3.
Without the harness, Claude Opus 5 scored 30% on ARC-AGI-3, which was the top result among all models tested.
ARC-AGI-3 is a benchmark of 2D games with no instructions; the model must work out how to play and win, and a 100% score means the model can beat the games as well as humans.
Adel El Hallack, vice president of product in Nvidia's AI unit, said the world interprets an agent almost as an API of the model, but an agent is the model plus the scaffolding around it (the harness), the runtime, and the associated skills and libraries it is given access to.
El Hallack said the more interesting part was introducing a supervising agent alongside the main working agent, which nudges the agent when it goes off direction, starts exploring a path that leads to a dead end, or re-explores a path it had previously trod.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-publisher vendor result, no primary paper or replication
The core 30%-to-100% finding rests on one TechCrunch article relaying Nvidia's own research and an Nvidia executive's quotes, with no paper link, evaluation protocol, run variance, per-model score table, or independent replication. Third-party corroboration exists only for the weaker surrounding thesis that harness choice matters (OpenAI's 3x tuning result, Microsoft's 19-model study, Databricks' cost finding), all cited second-hand.
Research-stage harness; ecosystem still on single-layer scaffolds
AVO is explicitly not a product and there is no disclosed production deployment, customer, or download signal. The only usage datapoint runs the other way: most agent users today rely on one harness layer such as Claude Code, Codex or Hermes. Adoption evidence is therefore limited to benchmark runs plus Nvidia's Nemo components being partly open.
Overstated: 'real hero' framing on one unreplicated 100% run
A perfect score on one interactive benchmark, produced by the vendor of the surrounding stack and framed as the harness being 'the real hero', outruns the supplied evidence. There is no ablation separating memory handling from supervision, no overhead or cost accounting for the extra agent layer, and no third party has reproduced 100%. The directional claim that harness matters is reasonably supported; the magnitude and generality claims are not.
Vendor-sourced with clear open-stack commercial interest
The findings and the interpretation both come from Nvidia, whose product VP uses them to argue that open harnesses, runtimes and infrastructure — the layers Nvidia sells and seeds via Nemo — put users in control, while pointedly citing OpenAI's low scores and training slowdown. Databricks' cost commentary is likewise supplied by its CEO. No countervailing voices from the labs characterised appear in the record.
Moderate-low: direction credible, magnitude unverified
Confidence is limited by a single-source cluster and vendor-originated headline numbers, but lifted by three independently reported studies pointing the same direction (harness tuning tripled OpenAI scores, all 19 Microsoft-tested models failed long-horizon editing, harness choice can 2x cost) and by an unambiguous, quotable mechanism description.
build
NVIDIA's safety teams put the agent security boundary in the runtime, not the model1 distinct publisher
build
Nvidia says the harness, not the model, took Claude Opus 5 from 30.2% to 100% on ARC-AGI-34 distinct publishers
build
Gartner: agent inference cost rises 5x by 2028, so budget per workflow, not per model1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026