Product1 publisher3 min readPublished
OpenAI's AGI-era claim rests on a benchmark score its own harness produced
Greg Brockman told the GPT-6 Astra launch that the AGI era had begun. The benchmark behind the claim returns a different score depending on whose testing software runs it, and the researcher who built it says it does not prove AGI.
The Product Desk · Product desk

What happened
- OpenAI president Greg Brockman told the launch event for the company's new GPT-6 Astra model earlier this month that the "AGI era" had begun.
- Astra scored 99.9 percent on the ARC-AGI-3 benchmark through OpenAI's own harness and 62.7 percent through the standard setup the benchmark supplies to give models a uniform interface.
- Astra set a record 169 on Epoch AI's Epoch Capabilities Index against a previous high of 163, and Epoch's analysis found the jump consistent with the existing trajectory of AI progress.
- Anthropic researcher Jacob Coxon resigned last week, saying he believes AI labs currently cannot mitigate the risks of increasingly intelligent and autonomous systems.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint A score that moves 37.2 points with the testing software cannot settle a procurement decision on its own, so the eval has to be rerun on the interface the buyer will actually ship.
- decision Teams writing model eval docs now have to record the harness beside the number, or accept a figure nobody in the building can reproduce.
- precedent With no agreed threshold, "AGI era" is available to any lab on any launch day, and the next one has been shown the price of using it: a week of researchers disputing the label.
- exposure Anyone who repeated the AGI framing internally to unlock budget owns the gap when the model performs in daily use much like the one it replaced.
A harness is the software layer that lets a model interact with a benchmark [6]. It decides what the model sees each turn, how many turns it gets, how its output is read back. Swap the harness and the score moves. The gap between Astra's two ARC-AGI-3 runs is 37.2 points [8], on the same model and the same games, and the lower figure came from the setup OpenAI did not build.
Buyers should care about the lower figure, because the wrapper your team writes will not be OpenAI's launch harness either. ARC-AGI-3 was built by Francois Chollet to resist training for the test, dropping a model into an unfamiliar video game and asking it to figure out the rules and win efficiently [7]. Chollet told Fast Company by email that "solving the benchmark is a strong sign of progress (as prior systems did not exhibit these attributes), but it is not proof of AGI, and that was never the point" [9]. He also listed what the games leave out: "The real world features much longer time horizons for continual learning compared to ARC 3 games (decades vs minutes), much larger world modeling complexity, much greater goal ambiguity, more greater exploration spaces, etc." [10]
The composite score points the same way. Epoch AI's index, which pools results from many benchmarks, moved six points to Astra's 169, a gain of about 3.7 percent over the prior record [15]. Gary Marcus, the NYU professor, wrote on his Substack earlier this month that a model that truly represented AGI should be outperforming rivals instead of matching them in everyday use [12]. "We are nowhere near AGI," Marcus said in a message to Fast Company. "That's just marketing by people who either don't know the original definitions or are deliberately lowering the bar." [11]
Andy Konwinski, who cofounded Databricks, Perplexity and Laude, described the gap in terms of work rather than scores. "These systems can't yet think on their own for long," he told Fast Company. "They're really good at coding, but most of the world's value doesn't come from software engineers." [16]
Nvidia's Jensen Huang endorsed the declaration, tweeting "AGI has arrived" [2]. Fast Company compared the whole exchange to the race among cellular carriers to put the next "G" on their networks before the underlying technology satisfied the technical standard [18]. The comparison works because there is no universally accepted threshold for AGI, and Chollet, Marcus and Konwinski all dispute that current models have crossed whatever it is [3].
For the person who has to put a model into a product this quarter, two columns are enough. Column one: the vendor's number and the harness that produced it. Column two: the number your team reproduced, on your own tasks, through your own interface, over the time horizon your users actually work in. The pitch will be built out of column one, and column two is the one you will be defending after launch. On ARC-AGI-3, the closest public stand-in for column two is the standard-interface run at 62.7 percent [5].
What to watch
- Whether anyone outside OpenAI reproduces the 99.9 percent ARC-AGI-3 run.
- Whether Chollet or OpenAI publishes what the OpenAI harness does that the standard interface does not.
- Whether Epoch AI's next index update shows a model departing from the trajectory Astra tracked.