ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness and 99.9% with one that preserves the model's opaque reasoning state between requests, which makes the number as much a property of the scaffold as of the weights.
Publishers:arcprize.org · superpowerdaily.com Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+30
- Incentives55
- Confidence65
build2 publishersConfirmed Artificial Analysis scores the same model level with its predecessor, and OpenAI charges two and a half times as much per token, so the ranking you inherit is a claim about a test mix that is not yours.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+35
- Incentives60
- Confidence58
build2 publishersConfirmed Vals AI's 141-hour Minecraft run ended with OpenAI's newest model farming potatoes for hours after losing its end-game loot and its spawn point in one explosion, and the recovery policy it came away with was a note to self.
Reality
- Evidence45
- Adoption30
- Hype gap+32
- Incentives68
- Confidence55
The capacity line carries no date and no named owner, so it cannot go into anyone's plan. The number an evaluation team can use is the 37.2 points between Astra's score in OpenAI's harness and its score in ARC Prize's.
Reality
- Evidence58
- Adoption44
- Hype gap+58
- Incentives86
- Confidence56
The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.
Perspective Coverage
3 publishers
- Builder
- Builder 27%
- Operator
- Operator 37%
- Investor
- Investor 36%
Reality
- Evidence66
- Adoption32
- Hype gap+61
- Incentives79
- Confidence71
ARC Prize ran GPT-6 Astra under two scaffolds and published both, 62.7% through its own minimal interface and 99.9% through OpenAI's Provider Adapter, which cost less. An eval that leaves the harness loose is scoring plumbing.
Reality
- Evidence68
- Adoption55
- Hype gap+64
- Incentives72
- Confidence62
The bill borrows the sentencing range used for unlawful nuclear weapons work. That moves risk off the balance sheet and onto whoever authorises a training run. It also landed on a day two harnesses scored the same model 35.9 points apart.
Reality
- Evidence26
- Adoption20
- Hype gap+44
- Incentives76
- Confidence30
ARC Prize put GPT-6 Astra at 62.7% against OpenAI's 99.9%, and the third-party composite index has it 0.3 points above the model it replaces, which leaves the 20% safety compute overhead as the clearest number in the launch.
Reality
- Evidence57
- Adoption42
- Hype gap+58
- Incentives74
- Confidence54
build1 publisherOne report Mithil Vakde's from-scratch transformer landed one point behind TRM on the public eval. His own ablations put 20 of those 44 points on two representation choices rather than on any amount of compute.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives72
- Confidence41
build1 publisherOne report One model was verified at 30.16 in the official harness and reported at 100.00 in NVIDIA's. Microsoft's Agent Lightning now trains the harness into the weights.
Reality
- Evidence54
- Adoption58
- Hype gap+28
- Incentives66
- Confidence48
build4 publishersConfirmed A five-person Nvidia team published the architecture behind its AVO agent alongside a perfect public-set score. The model did not change. The system around it did.
Perspective Coverage
4 publishers
- Builder
- Builder 50%
- Operator
- Operator 29%
- Investor
- Investor 21%
Reality
- Evidence58
- Adoption20
- Hype gap+34
- Incentives82
- Confidence72