Skip to content

ARC Prize

Non-profit organization that runs the ARC-AGI benchmark series, testing AI systems on novel reasoning puzzles to gauge progress toward general intelligence.

Known aliases

  • ARC Prize 2024 Technical Report
  • ARC Prize Foundation
  • arcprize.org

Current stories

science2 publishersConfirmed

Two harnesses put the same model 37 points apart on ARC-AGI-3

ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness and 99.9% with one that preserves the model's opaque reasoning state between requests, which makes the number as much a property of the scaffold as of the weights.

Publishers:arcprize.orgsuperpowerdaily.com

Reality

Evidence62
Adoption
Insufficient
Hype gap+30
Incentives55
Confidence65
leadership3 publishersConfirmed

ARC Prize puts Astra 37 points below the score OpenAI led with

The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.

Perspective Coverage

3 publishers
Builder
Builder 27%
Operator
Operator 37%
Investor
Investor 36%

Reality

Evidence66
Adoption32
Hype gap+61
Incentives79
Confidence71
invest1 publisherOne report

Astra's 99.9% holds up only on the harness OpenAI ran itself

ARC Prize put GPT-6 Astra at 62.7% against OpenAI's 99.9%, and the third-party composite index has it 0.3 points above the model it replaces, which leaves the 20% safety compute overhead as the clearest number in the launch.

Publishers:techtimes.com

Reality

Evidence57
Adoption42
Hype gap+58
Incentives74
Confidence54
build4 publishersConfirmed

Nvidia says the harness, not the model, took Claude Opus 5 from 30.2% to 100% on ARC-AGI-3

A five-person Nvidia team published the architecture behind its AVO agent alongside a perfect public-set score. The model did not change. The system around it did.

Perspective Coverage

4 publishers
Builder
Builder 50%
Operator
Operator 29%
Investor
Investor 21%

Reality

Evidence58
Adoption20
Hype gap+34
Incentives82
Confidence72