Epoch AI, which benchmarks AI systems, had mathematicians nominate 50 high-stakes open problems chosen so a computer can check any proposed solution. The rule lets research mathematics be graded as a test of AI, but only problems whose answers a machine can confirm get in.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence40
Thieves broke into two PlusAI trailers taken from Fremont, one Nvidia-branded, and found 40,000 pounds of test sand. Overhaul's tally of more than $150 million in AI cargo losses is small next to AI spending, but single thefts run into the millions, and that cost lands on whoever ships each load of chips.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives65
- Confidence45
Andreessen Horowitz puts hyperscaler capex near $780 billion this year and tech at roughly 55% of US capital spending. Its charts also show free cash flow being squeezed, and the hyperscalers are relying more on debt to fund new chip capacity.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives75
- Confidence60
Epoch AI puts the price of answering a GPQA Diamond question at a fixed accuracy bar falling about 13 times a year, faster than compute under Moore's Law. Buyers get that discount only by moving to newer models, so the part of a stack that has to switch cheaply is the evaluation that qualifies each one.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence45
US companies are buying the cheapest AI model that can finish a job, the FT reports, with Ramp data pointing to a 41% fall in effective token prices. Investors now have to value AI vendors on cost per completed task, a figure that shifts with the workload.
Reality
- Evidence42
- Adoption50
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
The build is now the largest driver of U.S. private investment growth, and 71% of Americans polled want none of it nearby, which makes entitlement risk rather than capital the thing infrastructure investors have to price.
Perspective Coverage
4 publishers
- Builder
- Builder 8%
- Operator
- Operator 46%
- Investor
- Investor 46%
Reality
- Evidence58
- Adoption62
- Hype gap+45
- Incentives64
- Confidence60
Artificial Analysis scores the same model level with its predecessor, and OpenAI charges two and a half times as much per token, so the ranking you inherit is a claim about a test mix that is not yours.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+35
- Incentives60
- Confidence58
OpenAI reports 3.1 agent-workdays of agent runtime for every human workday, then spends much of the same report explaining why research did not get 3.1 times faster. The residue lands on review, compute allocation and deciding what to run.
Reality
- Evidence55
- Adoption70
- Hype gap+25
- Incentives60
- Confidence60
Anthropic has put a number on how much of its own AI research Claude now runs. The number comes out of a pipeline in which a Claude agent catalogued the tasks and a separate Claude judge scored them.
Reality
- Evidence35
- Adoption55
- Hype gap+20
- Incentives70
- Confidence40
The lab says Claude leads 26% of its model research and collaborates on about 90%, with both tiers defined by how closely a human directs each task, and it wants rival labs publishing the same measure.
Perspective Coverage
6 publishers
- Builder
- Builder 39%
- Operator
- Operator 33%
- Investor
- Investor 28%
Reality
- Evidence50
- Adoption65
- Hype gap+30
- Incentives70
- Confidence60
Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.
Reality
- Evidence65
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence62
Anthropic says Claude now leads 26% of its model research and development under human supervision, and it wants other labs to publish comparable figures. The labs count different things.
Perspective Coverage
4 publishers
- Builder
- Builder 23%
- Operator
- Operator 41%
- Investor
- Investor 36%
Reality
- Evidence56
- Adoption52
- Hype gap+24
- Incentives72
- Confidence66
Dario Amodei's essay asks frontier labs to pace capability gains and asks for antitrust room to coordinate on safety. Europe's challengers, holding 11% of global AI venture funding, say the leaders wrote it for themselves.
Reality
- Evidence45
- Adoption22
- Hype gap+22
- Incentives80
- Confidence55
Anthropic says Claude leads 26 percent of its model research and development and takes part in more than 90 percent of it. The company wants rival labs to publish the same index so the numbers can be compared over time.
Publishers:ca.finance.yahoo.com · engadget.com · fastcompany.com Perspective Coverage
3 publishers
- Builder
- Builder 30%
- Operator
- Operator 42%
- Investor
- Investor 28%
Reality
- Evidence46
- Adoption44
- Hype gap+24
- Incentives74
- Confidence62
Red Hat clocks the same 20-call agent task at roughly 45 seconds on a slow backend and about 13 on a fast one. Model choice for agents is turning into a per-call latency budget, with capability as one input.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives50
- Confidence55
The lab published a prototype index of how much of its own AI research Claude now leads, alongside a promise to embed outside evaluators with access comparable to what its internal risk teams get. The figures are dated August 2026.
Reality
- Evidence38
- Adoption32
- Hype gap+22
- Incentives74
- Confidence46
A C4ADS investigation maps three routes for restricted accelerators into China, and almost all of the value it counted sits with a single importer whose ownership is still unresolved.
Reality
- Evidence52
- Adoption45
- Hype gap+28
- Incentives62
- Confidence50
Vinod Bijlani argues in Forbes that a token price measures only what an application consumed, while the idle accelerators and rising power draw behind it never reach the invoice a CIO reviews. He works for HPE.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+22
- Incentives76
- Confidence52
Astra is rated cheaper per task on many benchmarks. After putting it in front of roughly 3,500 engineers, Databricks reports about 60 percent more coding spend.
Reality
- Evidence38
- Adoption62
- Hype gap+30
- Incentives60
- Confidence45
Greg Brockman told the GPT-6 Astra launch that the AGI era had begun. The benchmark behind the claim returns a different score depending on whose testing software runs it, and the researcher who built it says it does not prove AGI.
Reality
- Evidence55
- Adoption25
- Hype gap+55
- Incentives80
- Confidence58
Earlier coverage
- Nomic's CEO alleges startups resell frontier traces generated on lab credits
Build · September 14, 2026 · 1 publisher
- Micro1 outbids Google by 25% for Spirit Airlines' decades of records
Invest · September 14, 2026 · 1 publisher
- Huawei's 750,000 Ascend chips add up to under a twenty-fifth of Nvidia's 2026 compute
Invest · September 9, 2026 · 1 publisher
- AI text detectors are good enough to deploy. The appeals process is what nobody has written.
Science · August 25, 2026 · 1 publisher
- Sutton calls synthetic data 'a big mistake': every simulator is a lossy copy of a bigger world
Product · August 19, 2026 · 1 publisher
- Re-baseline AI procurement on cost per completed task, not dollars per million tokens
Leadership · August 18, 2026 · 1 publisher
- Your Training Data Now Has Counterparties: The Case For Treating Corpora As Procurement
Build · August 17, 2026 · 1 publisher
- Microsoft has 2.2m AI chips installed. Its own capacity claims imply up to 6.4m
Leadership · August 17, 2026 · 1 publisher
- Artificial Analysis moves eval onto your data, and turns model choice into procurement
Build · August 15, 2026 · 1 publisher