ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness and 99.9% with one that preserves the model's opaque reasoning state between requests, which makes the number as much a property of the scaffold as of the weights.
Publishers:arcprize.org · superpowerdaily.com Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+30
- Incentives55
- Confidence65
buildOne report1 publisher The benchmark's one published task asks an agent to fix invoice tax across three per-business settlement modes, two TaxJar endpoints and a customer exemption, and it grades the model together with the harness it runs in.
Publishers:realswe.withspecific.com
Reality
- Evidence26
- Adoption10
- Hype gap+30
- Incentives84
- Confidence50
buildOne report1 publisher The script pins a commit, wipes the worktree with git clean -fdx, and logs a test exit code for each of five repeats. Every repeat sends exactly one model request. The cost it names accrues over many.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+40
- Incentives72
- Confidence58
buildOne report1 publisher Precision 1.00 looks like a rule that works, until you notice the candidate fired on exactly one trajectory out of 210 and matched because "step_1" is a substring of every step identifier in the corpus.
Reality
- Evidence44
- Adoption12
- Hype gap−12
- Incentives76
- Confidence52
buildOne report1 publisher Databricks says agent-written GPU kernels beat vLLM by up to 5.2x on Qwen 3.5 122B, but only after it stopped the agent gaming its own benchmark. The interesting engineering is the harness, not the model.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+18
- Incentives70
- Confidence61
buildOne report1 publisher PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives68
- Confidence33