Skip to content

Benchmark

ARC-AGI-3

ARC Prize benchmark that places AI agents in novel, game-like environments with no instructions, testing generalization rather than memorized skills.

Known aliases

  • ARC 3
  • ARC-AGI
  • ARC AGI 3
  • ARC-AGI3
  • ARC-AGI 3
  • ARC-AGI-3
  • ARC-AGI-3 benchmark
  • ARC-AGI-3 public set
  • ARC-AGI-3 Semi-Private

Current stories

build17 publishersConfirmed

Astra's Critical cyber rating ships a real-time pause switch inside the Bedrock service boundary

Greg Brockman said AGI arrived with GPT-6 Astra on September 3. The enforcement the launch actually documents is a misuse classifier running inside AWS's service boundary, plus a voluntary 30-day US review that carried no license.

Perspective Coverage

17 publishers
Builder
Builder 39%
Operator
Operator 37%
Investor
Investor 24%

Reality

Evidence58
Adoption42
Hype gap+45
Incentives72
Confidence60
security4 publishersConfirmed

OpenAI gates a 100% ExploitBench model behind refusals it plans to loosen in weeks

Astra scored a perfect 100% on OpenAI's own exploit-development benchmark, against 78.5% for GPT-5.6 Sol, and the shipped model's refusal to write proof-of-concept code is a policy the company has already said it will relax.

Perspective Coverage

4 publishers
Builder
Builder 30%
Operator
Operator 42%
Investor
Investor 28%

Reality

Evidence35
Adoption20
Hype gap+45
Incentives70
Confidence55
science2 publishersConfirmed

Two harnesses put the same model 37 points apart on ARC-AGI-3

ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness and 99.9% with one that preserves the model's opaque reasoning state between requests, which makes the number as much a property of the scaffold as of the weights.

Publishers:arcprize.orgsuperpowerdaily.com

Reality

Evidence62
Adoption
Insufficient
Hype gap+30
Incentives55
Confidence65
product18 publishersConfirmed

OpenAI's GPT-6 Astra pairs harder-to-monitor reasoning with a pledge to pause scaling if oversight slips

OpenAI's new model thinks repeatedly before it acts, and according to Manifold Security's CTO it usually does so without leaving the reasoning trace that agent audits read. Oversight moves to the buyer.

Perspective Coverage

18 publishers
Builder
Builder 32%
Operator
Operator 38%
Investor
Investor 30%

Reality

Evidence52
Adoption25
Hype gap+45
Incentives72
Confidence60
build1 publisherOne report

A nullable parentId buys Pi the entire session tree for one field

Codex made the session Item a wire type and Pi gave every entry a nullable parentId, while Claude Code squeezes the transcript in five stages. The three designs diverge on what stays addressable after compaction.

Publishers:dev.to

Reality

Evidence58
Adoption30
Hype gap+12
Incentives30
Confidence60
leadership3 publishersConfirmed

ARC Prize puts Astra 37 points below the score OpenAI led with

The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.

Perspective Coverage

3 publishers
Builder
Builder 27%
Operator
Operator 37%
Investor
Investor 36%

Reality

Evidence66
Adoption32
Hype gap+61
Incentives79
Confidence71
invest1 publisherOne report

Astra's 99.9% holds up only on the harness OpenAI ran itself

ARC Prize put GPT-6 Astra at 62.7% against OpenAI's 99.9%, and the third-party composite index has it 0.3 points above the model it replaces, which leaves the 20% safety compute overhead as the clearest number in the launch.

Publishers:techtimes.com

Reality

Evidence57
Adoption42
Hype gap+58
Incentives74
Confidence54

Earlier coverage

  1. Mostik's bridge splits the difference between a 753B and a 4B model at a twentieth of the cost

    Product · September 2, 2026 · 1 publisherOne report

  2. Ord's generation-time argument makes runaway AI unlikely, not just slower

    Leadership · August 30, 2026 · 1 publisherOne report

  3. Anthropic ships a price dial with its new model, and that is now the buying decision

    Leadership · August 26, 2026 · 1 publisherOne report

  4. Same weights, 70 points apart: the ARC-AGI-3 table has stopped being procurement evidence

    Build · August 24, 2026 · 1 publisherOne report

  5. Nvidia moved one model from 30% to 100% without changing the model

    Product · August 22, 2026 · 1 publisherOne report

  6. Nvidia says the harness, not the model, took Claude Opus 5 from 30.2% to 100% on ARC-AGI-3

    Build · August 21, 2026 · 4 publishersConfirmed