Sitefire's SlopShape spots AI-written company blog posts at 97.0% macro-F1 from page structure alone, and at 96.1% after each model rewords its own output. Paraphrasing generated copy hides little from it, though the test reworded text without reordering any page's sections.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives45
- Confidence35
OpenAI's deprecations page lists about fifty model snapshots, GPT-4 among them, with shutdown dates between September 24 and December 11, 2026. How each team's code fails on those dates depends on whether it pinned a snapshot or called an alias.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap0
- Incentives50
- Confidence55
Novices using InstructMesh, a tool from MIT CSAIL, Google and Northeastern, found and fixed flaws in AI-generated 3D models about 90 percent of the time. An expert graded the edited designs, so how the printed objects perform is still an open question.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence45
Instinct raised $1 billion at a $10 billion valuation, four times the price TechCrunch reported for its August round. Its agent works through users' own screens and accounts, so permission and recovery are Instinct's to engineer before any app maker is asked.
Perspective Coverage
3 publishers
- Builder
- Builder 23%
- Operator
- Operator 27%
- Investor
- Investor 50%
Reality
- Evidence55
- Adoption20
- Hype gap+45
- Incentives60
- Confidence60
The top rung of OpenAI's Preparedness Framework has now been reached by OpenAI, on a model it has not shipped, which moves AI-assisted exploitation out of argument and into a named vendor's published paperwork.
Perspective Coverage
5 publishers
- Builder
- Builder 28%
- Operator
- Operator 42%
- Investor
- Investor 30%
Reality
- Evidence35
- Adoption3
- Hype gap+30
- Incentives70
- Confidence55
Per-unit intelligence has kept getting cheaper since GPT-4 shipped in 2023, and enterprise AI bills have kept climbing anyway. Two McKinsey senior partners spent a Tuesday session on why consumption and vendor margins explain it.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence50
OpenAI's own deprecations page lists four cutoffs between 24 September and 23 October 2026. Finding out whether they touch your code takes one look at the Usage dashboard, in an account plenty of buyers cannot sign into.
Reality
- Evidence66
- Adoption62
- Hype gap+16
- Incentives58
- Confidence64
The 5% comes from one dev.to post's worked example, and it holds only if an attacker's hit rate against your login is near 0.3% and a stolen account resells near $20. At a 0.03% hit rate the same bill is half of revenue.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+52
- Incentives58
- Confidence30
The failure modes in a dev.to writeup on shipping LLM apps are each ordinary engineering work with an owner. Its cost example implies a per-call price 3.6 to 9 times below the demo figure it starts from.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives25
- Confidence35
Anil Madhavapeddy patched a path traversal bug in OCaml's cohttp and found probes for it in his logs ten minutes after opening the fix PR. His own agent had already built the exploit from a bug-class hint.
Publishers:anil.recoil.org
Reality
- Evidence52
- Adoption44
- Hype gap+14
- Incentives55
- Confidence57
A 91-run simulation put 100 LLM agents on Pokhara Lakeside geography for 26 simulated weeks and reports rigid prices and stuck wages. Deleting the agents' memory changed nothing the team could measure.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+40
- Incentives62
- Confidence30
Lu, Chen and Wu store an agent's procedure as typed nodes and edges, and an LLM refiner proposes changes by comparing failed runs with successful ones. Keeping an edit requires a validation set, so adoption starts with scored trajectory logs.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+46
- Incentives38
- Confidence29
Y Combinator's chief executive would leave distillation alone and have regulators police the gap between open weight and frontier pricing, while the NSA, CISA and FBI treat the copying as a cyber security matter.
Reality
- Evidence52
- Adoption30
- Hype gap+18
- Incentives78
- Confidence58
The paper reports matching GPT-4 with up to 98 percent less spend. The saving rests on a 150x spread in prompt prices between providers and on a triage rule fitted separately per dataset and task.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives60
- Confidence55
A systematic review grades LLM work in building automation by deployment readiness, and the pattern that clears the bar translates raw point names into a canonical schema while a human signs off on the table.
Reality
- Evidence38
- Adoption8
- Hype gap+10
- Incentives
- Insufficient
- Confidence42
The fight over data center water is argued in milliliters per query, while the most cited US projection says the deciding variables are which plant fills the power contract and which basin the campus sits in.
Reality
- Evidence58
- Adoption40
- Hype gap+12
- Incentives45
- Confidence55
A computer engineer writing in Fast Company says dermatology models latch onto background skin tone instead of the lesion, which turns training-set composition into a spec question for anyone shipping skin imaging to patients.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+18
- Incentives40
- Confidence45
A team from King's College London and UCL puts the harm in ordinary product behaviour, sycophancy trained in through RLHF plus a session whose state the user writes, and argues mitigation should not wait for a label.
Reality
- Evidence32
- Adoption20
- Hype gap+30
- Incentives45
- Confidence36
SWE-Gate scores agent patches at two gates instead of one. The third that clear tests but fail mined review rules are measured against a pipeline that runs tests and nothing else, while most teams run a pipeline with additional checks beyond that.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+45
- Incentives25
- Confidence35
The Information says Astra cycles its thinking through internal layers instead of writing it out, and OpenAI's chief scientist has answered the report without confirming the architecture, which leaves anyone monitoring reasoning guessing.
Reality
- Evidence44
- Adoption17
- Hype gap+31
- Incentives66
- Confidence52
Earlier coverage
- Bletchley's insider-trading demo warned of AI deception - now incidents are surging
Leadership · September 1, 2026 · 1 publisher
- AI is absorbing the grunt work that trained the next generation of seniors
Leadership · September 1, 2026 · 1 publisher
- Liner's $36.1M bet is that the footnote, not the model, is the enterprise product
Product · August 25, 2026 · 1 publisher
- Encrypted reasoning that a cheaper sibling model can open is not protected IP
Build · August 24, 2026 · 1 publisher
- GitHub's Java agent runtime ships as a Maven dependency, and the tool schema comes from reflection
Build · August 22, 2026 · 1 publisher
- JFrog measured 847 log lines to find 9, and that ratio is now a budget line
Build · August 20, 2026 · 1 publisher
- GPT-4 as a scaling instrument: it matched human emotion maps, then went past them
Science · August 19, 2026 · 1 publisher
- Re-baseline AI procurement on cost per completed task, not dollars per million tokens
Leadership · August 18, 2026 · 1 publisher
- Nature Perspective: patching one fact into a model leaves the reasoning around it broken
Science · August 16, 2026 · 1 publisher