Build4 publishersReports disagree3 min readPublished
Nvidia says the harness, not the model, took Claude Opus 5 from 30.2% to 100% on ARC-AGI-3
A five-person Nvidia team published the architecture behind its AVO agent alongside a perfect public-set score. The model did not change. The system around it did.
The Engineer · Build desk

What happened
- Nvidia first introduced its Agentic Variation Operators (AVO) general-purpose coding agent system in late March 2026, and has now published the architecture and system-level mechanisms behind it, applied to the ARC-AGI-3 benchmark.
- In a team blog released on Friday, a five-person team of Nvidia software engineers, machine learning specialists and AI research interns, headed by principal engineer Terry Chen, described how AVO elevated Claude Opus 5's ARC-AGI-3 performance.
- ARC Prize, a non-profit AI research and benchmarking body, used an analysis post this July to report a 30.2% score for Claude Opus 5 at high reasoning effort on the public set of environments and tasks in ARC-AGI-3.
- AVO achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels.
- ARC-AGI-3 uses Relative Human Action Efficiency (RHAE), a metric combining task completion with per-level action efficiency relative to first-time human baselines, aggregated across different levels and environments.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Nvidia has published the architecture of Agentic Variation Operators (AVO), the general-purpose coding agent system it introduced in late March 2026, and applied it to ARC-AGI-3 [s1 c1]. In a blog posted Friday by a five-person team of engineers, machine learning specialists and research interns led by principal engineer Terry Chen, the company reports that Claude Opus 5 went from a 30.2% baseline on the benchmark to a 100.00 RHAE score, with the harness as the only thing that changed [10][2][3].
The numbers first. ARC Prize, the non-profit that maintains the benchmark, reported 30.2% for Claude Opus 5 at high reasoning effort on the ARC-AGI-3 public set in an analysis post this July [2]. Nvidia reports 100.00 RHAE across all 25 environments in that public set, with all 183 levels completed [3]. That is a gap of 69.8 points, or roughly 3.3 times the baseline [22]. RHAE combines task completion with per-level action efficiency measured against first-time human baselines, aggregated across levels and environments [12], so it is not a simple pass rate and a perfect score means the agent was also action-efficient.
What makes this worth reading rather than filing is the claim about where the capability came from. Nvidia's framing is that the harness determines how a model receives context, uses tools, maintains state, responds to feedback, recovers from failure and sustains progress on long-running work [21]. The team writes that the result shows system design, not model capability alone, can unlock frontier-level long-horizon performance, and that "model capability matters enormously, but the surrounding system determines how effectively that capability can be converted into sustained autonomous progress" [20][19].
The transfer story is the part operators should look at. AVO was built for software engineering and GPU-kernel optimization, where it replaces the predefined variation step of conventional evolutionary search with an agent that decides what to inspect, change, test and commit [6]. For ARC-AGI-3, Nvidia says it connected the same agent to a different task interface, changing only the environment-specific tools and the evaluation [5]. ARC-AGI-3 drops agents into unfamiliar environments with no instructions, stated rules or stated goals [11], which on the surface has nothing to do with compilers and throughput, but Nvidia argues the loop is the same: build hypotheses from incomplete evidence, act through an interface, observe consequences, preserve state, revise, recover from wrong assumptions, keep going [13].
Two mechanisms carry that loop past a single context window: persistent memory and supervision [7]. Memory holds prior implementations, evaluation results, compiler and profiler output and accumulated reasoning so the agent resumes instead of rebuilding the search [9]. A supervisor watches the trajectory for stagnation and repeated unproductive cycles and can redirect the main agent [4]. In the kernel work, AVO ran continuously for seven days, explored more than 500 optimization directions and committed 40 kernel versions, with the supervisor stepping in when the search plateaued [15][8]. The resulting multihead attention kernels beat cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% on DGX B200 across the configurations evaluated [16], and the agent then adapted the kernel to grouped-query attention in about 30 minutes [17].
Two things to watch. The 100.00 is a public-set result [3]; ARC Prize's own comparative figures for Claude Opus 5 include a semi-private ARC-AGI-2 score of 90.4% [18], and neither Nvidia post reports a semi-private ARC-AGI-3 number. And the baseline and the agent result were produced by different parties, ARC Prize and Nvidia respectively [2][3], which is the usual reason to wait for a third-party run before treating a 69.8-point delta as a property of the harness rather than of the setup.