Glow says AI coding agents asked for review screenshots put over 13,000 internal images from 300-plus organizations into public GitHub repositories. Most sat under developers' personal accounts, where the companies' security teams were not looking.
Perspective Coverage
5 publishers
- Builder
- Builder 37%
- Operator
- Operator 53%
- Investor
- Investor 10%
Reality
- Evidence55
- Adoption45
- Hype gap+15
- Incentives70
- Confidence60
AppZen says its ZenLM Plus finance models led seven frontier models on five of six expense-audit controls in a test the company ran itself. Until buyers rerun that test on their own expense policies, the scores describe AppZen's data and configuration.
Reality
- Evidence28
- Adoption10
- Hype gap+40
- Incentives78
- Confidence35
Huntress ran seven Rails tasks past Claude Fable 5.1 and watched it copy hand-rolled code out of its own repository. Changing the harness, in three cheap steps, took API recall from 48 percent to 100.
Reality
- Evidence58
- Adoption25
- Hype gap+15
- Incentives40
- Confidence55
A LessWrong team fixed an eight-character murder plot before any agent spoke, then graded monitors on what they reported and what they missed. The best one fully recovered under half the rubric's facts.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives30
- Confidence55
Jev returns typed answers with probability scores and no free-form text, and TypeSafe prices a decision at four hundredths of a cent. Its accuracy figures are scored against other models' probabilities.
Publishers:arize.com
Reality
- Evidence34
- Adoption22
- Hype gap+30
- Incentives68
- Confidence44
CommentBench splits human comments on AI-safety posts and drafts into target points, filters out the ones a model could not reach without extra context, and has Opus 5 judge which of the rest a model hit.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives55
- Confidence50
Vercel's September Production Index shows the gateway's average price per token down 23.2% in August, a third straight decline, and the median heavy-usage team paying 7.6% less, so most of the saving came from switching models.
Reality
- Evidence58
- Adoption80
- Hype gap+14
- Incentives70
- Confidence64
A LessWrong write-up planted invalidating flaws in ML experiment logs and asked models to write the conference abstract. A second model scored the disclosure on three levels, and the published example is one before-and-after pair.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence40
Nineteen days after the same model name produced repetition loops and phantom missing files, the served checkpoint refused 59.2% of concealed-hazard tasks. That makes third-party refusal testing a running job, not a procurement step.
Reality
- Evidence60
- Adoption28
- Hype gap+12
- Incentives72
- Confidence55
Johann Rehberger got code execution in up to 80% of attempts by making one tool fail and letting the agent improvise. Anthropic says Auto Mode is working as designed, which leaves the sandbox decision on your desk.
Reality
- Evidence63
- Adoption70
- Hype gap−8
- Incentives58
- Confidence66
Aikido spent 11.7 billion tokens rediscovering 32 fresh CVEs with ten models, three attempts each. The number that should move a scanning budget is the marginal cost of the second and third pass.
Reality
- Evidence44
- Adoption18
- Hype gap+32
- Incentives76
- Confidence52
SemiAnalysis puts 40-50% of 2027's new capacity under OpenAI and Anthropic contracts, with revenue per megawatt now well clear of deployment cost. That makes inference a residual market.
Reality
- Evidence20
- Adoption30
- Hype gap+55
- Incentives65
- Confidence26
The frontier price cut concedes that buyers now price intelligence per task. The discount expires in a quarter; Thomson Reuters' own domain model does not.
Reality
- Evidence42
- Adoption48
- Hype gap+30
- Incentives62
- Confidence44
Anthropic's promotional lift to weekly limits was set to end around 19 August. The invoice does not change, which is exactly why this repricing is hard to see.
Reality
- Evidence26
- Adoption24
- Hype gap+22
- Incentives52
- Confidence34
A production test across 15 models put seven of them inside a one-point spread on pass rate. On constrained payroll work, the price premium bought speed, not correctness.
Reality
- Evidence58
- Adoption34
- Hype gap+18
- Incentives55
- Confidence52