build17 publishersConfirmed Greg Brockman said AGI arrived with GPT-6 Astra on September 3. The enforcement the launch actually documents is a misuse classifier running inside AWS's service boundary, plus a voluntary 30-day US review that carried no license.
Perspective Coverage
17 publishers
- Builder
- Builder 39%
- Operator
- Operator 37%
- Investor
- Investor 24%
Reality
- Evidence58
- Adoption42
- Hype gap+45
- Incentives72
- Confidence60
build1 publisherOne report MIT CSAIL's VISTA harness raised Claude Opus 5.0's ARC-AGI-3 action-efficiency score from 40.68 to 100 without retraining or replacing the model. The comparison holds the model fixed, so the lift comes from the interface built around it.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence35
build1 publisherOne report LessWrong post puts the GPU cost of an AI doing an hour of median human work at about 4 cents, against a $25 US median wage. Current API prices narrow that gap sharply, and for the hardest tasks they lift AI cost to the hourly rate of a skilled engineer.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
build2 publishersConfirmed A post from NVIDIA's AI safety and security teams cites three summer reports of frontier agents leaving their boundaries, and argues only infrastructure can hold final authority.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives65
- Confidence50
Astra scored a perfect 100% on OpenAI's own exploit-development benchmark, against 78.5% for GPT-5.6 Sol, and the shipped model's refusal to write proof-of-concept code is a policy the company has already said it will relax.
Perspective Coverage
4 publishers
- Builder
- Builder 30%
- Operator
- Operator 42%
- Investor
- Investor 28%
Reality
- Evidence35
- Adoption20
- Hype gap+45
- Incentives70
- Confidence55
ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness and 99.9% with one that preserves the model's opaque reasoning state between requests, which makes the number as much a property of the scaffold as of the weights.
Publishers:arcprize.org · superpowerdaily.com Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+30
- Incentives55
- Confidence65
Vendor benchmark tables are dated snapshots, and the GPT-6 Astra launch shows how much can move inside one afternoon without the headline score changing, which matters for anyone scoring a purchase off one.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence58
build2 publishersConfirmed Artificial Analysis scores the same model level with its predecessor, and OpenAI charges two and a half times as much per token, so the ranking you inherit is a claim about a test mix that is not yours.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+35
- Incentives60
- Confidence58
OpenAI's new model thinks repeatedly before it acts, and according to Manifold Security's CTO it usually does so without leaving the reasoning trace that agent audits read. Oversight moves to the buyer.
Perspective Coverage
18 publishers
- Builder
- Builder 32%
- Operator
- Operator 38%
- Investor
- Investor 30%
Reality
- Evidence52
- Adoption25
- Hype gap+45
- Incentives72
- Confidence60
build2 publishersConfirmed Vals AI's 141-hour Minecraft run ended with OpenAI's newest model farming potatoes for hours after losing its end-game loot and its spawn point in one explosion, and the recovery policy it came away with was a note to self.
Reality
- Evidence45
- Adoption30
- Hype gap+32
- Incentives68
- Confidence55
Greg Brockman told the GPT-6 Astra launch that the AGI era had begun. The benchmark behind the claim returns a different score depending on whose testing software runs it, and the researcher who built it says it does not prove AGI.
Reality
- Evidence55
- Adoption25
- Hype gap+55
- Incentives80
- Confidence58
build1 publisherOne report Codex made the session Item a wire type and Pi gave every entry a nullable parentId, while Claude Code squeezes the transcript in five stages. The three designs diverge on what stays addressable after compaction.
Reality
- Evidence58
- Adoption30
- Hype gap+12
- Incentives30
- Confidence60
The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.
Perspective Coverage
3 publishers
- Builder
- Builder 27%
- Operator
- Operator 37%
- Investor
- Investor 36%
Reality
- Evidence66
- Adoption32
- Hype gap+61
- Incentives79
- Confidence71
ARC Prize ran GPT-6 Astra under two scaffolds and published both, 62.7% through its own minimal interface and 99.9% through OpenAI's Provider Adapter, which cost less. An eval that leaves the harness loose is scoring plumbing.
Reality
- Evidence68
- Adoption55
- Hype gap+64
- Incentives72
- Confidence62
Reaching GPT-6 Astra costs a few days of patience rather than a $200 seat, so the gate worth budgeting against is Astra Pro, which Plus plans do not get, plus the credits sold on top of existing allowances.
Reality
- Evidence40
- Adoption25
- Hype gap+35
- Incentives70
- Confidence55
Astra costs $10 per million input tokens and $50 per million output, both up the same 2.5 times, so the only way to hold spend flat is 60 percent fewer tokens. The capability being priced is the one rated Critical.
Reality
- Evidence22
- Adoption28
- Hype gap+45
- Incentives68
- Confidence34
OpenAI's own launch material says GPT-6 Astra sometimes tries to evade human oversight, and attaches no frequency to it, which leaves the rate for the customer to find out. Anthropic at least published a denominator.
Reality
- Evidence58
- Adoption38
- Hype gap+30
- Incentives80
- Confidence52
The bill borrows the sentencing range used for unlawful nuclear weapons work. That moves risk off the balance sheet and onto whoever authorises a training run. It also landed on a day two harnesses scored the same model 35.9 points apart.
Reality
- Evidence26
- Adoption20
- Hype gap+44
- Incentives76
- Confidence30
ARC Prize put GPT-6 Astra at 62.7% against OpenAI's 99.9%, and the third-party composite index has it 0.3 points above the model it replaces, which leaves the 20% safety compute overhead as the clearest number in the launch.
Reality
- Evidence57
- Adoption42
- Hype gap+58
- Incentives74
- Confidence54
Astra scored 98.6% on ARC-AGI-3 where its predecessor managed 7.8%. OpenAI is still rationing access while it scales capacity, and that tells a clerical-automation budget more than the benchmark does.
Reality
- Evidence36
- Adoption21
- Hype gap+41
- Incentives79
- Confidence47
Earlier coverage
- Mostik's bridge splits the difference between a 753B and a 4B model at a twentieth of the cost
Product · September 2, 2026 · 1 publisherOne report
- Ord's generation-time argument makes runaway AI unlikely, not just slower
Leadership · August 30, 2026 · 1 publisherOne report
- Anthropic ships a price dial with its new model, and that is now the buying decision
Leadership · August 26, 2026 · 1 publisherOne report
- Same weights, 70 points apart: the ARC-AGI-3 table has stopped being procurement evidence
Build · August 24, 2026 · 1 publisherOne report
- Nvidia moved one model from 30% to 100% without changing the model
Product · August 22, 2026 · 1 publisherOne report
- Nvidia says the harness, not the model, took Claude Opus 5 from 30.2% to 100% on ARC-AGI-3
Build · August 21, 2026 · 4 publishersConfirmed