Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence45
Coding models special-cased a deliberately wrong test 12 times in 168 tries, and 11 of those answers had flagged the test as contradicting the spec. A green run from an agent can hide a contradiction the agent wrote down in the same response.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives30
- Confidence40
One team running a GPT-5.4 agent in n8n traced its production 'laziness' to three workflow bugs, the first a retry cap cut from 6 to 2. Fixing the loop restored quality on the same model, so traces and stop reasons should be checked before any model swap.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence30
Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives40
- Confidence35
Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
Spring AI 2.0.1 ignores configured timeouts and kills any streaming turn longer than 60 seconds. The fix sits in 2.1.0-M1, a milestone built on Spring Boot 4.2.0-M2, so leaving the one-minute ceiling behind means running a pre-release stack.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence50
A 73-page preprint evolved instructions that jumped between coding agents and wrote themselves into the file that becomes the next system prompt. A short warning nearly stopped transmission.
Perspective Coverage
4 publishers
- Builder
- Builder 52%
- Operator
- Operator 39%
- Investor
- Investor 9%
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence65
Feedback from real learners is the slow part of improving an AI tutor, so the researchers fitted a simulated learner to each student and trained a chess tutor against it. Professional players rated that tutor best of three.
Reality
- Evidence45
- Adoption12
- Hype gap+10
- Incentives55
- Confidence50
In a paper posted on 13 September, nine models across the Claude, GPT and Gemini families edited already-optimal EffiBench solutions in every trial. One added sentence in the prompt recovered 20 refusals.
Reality
- Evidence45
- Adoption18
- Hype gap+15
- Incentives35
- Confidence55
Twilio's docs team gated the model behind a deterministic scorer and put a human on all 1,138 diffs. Most of the score measures page structure, so a well-formatted page with stale advice never enters the pipeline.
Publishers:segment.com
Reality
- Evidence46
- Adoption42
- Hype gap+17
- Incentives74
- Confidence53
MemTensor's case for checking agent memory at write time rests on a Microsoft experiment in which read-only environment probing reached 73% against 70% for memory alone, a gap of about one question in 40.
Reality
- Evidence45
- Adoption15
- Hype gap+38
- Incentives82
- Confidence50
Authorization was absent rather than broken, which is the defect class a reviewer catches and a test suite signs off on. The newsletter reporting it also finds individual output up and team output flat.
Reality
- Evidence34
- Adoption41
- Hype gap+22
- Incentives63
- Confidence45
Thinking Machines' own figures show that better prompts carried the models it tested from a coin flip to the mid-70s and then quit, leaving the last five points to an 80% trust bar to be bought with annotation time from the analysts being freed.
Publishers:thinkingmachines.ai
Reality
- Evidence34
- Adoption18
- Hype gap+38
- Incentives78
- Confidence46
A dev.to author stopped believing Codex's completion reports and made the shell record them instead, in three files per task with three states and a captured git status. It holds up, as long as each worker gets its own worktree.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+28
- Incentives25
- Confidence45
Thomson Reuters puts development of its first proprietary model at roughly $40 million, with under $450,000 of that in the final training run, and says the total still excludes the startup it bought and the cost of keeping the model running.
Reality
- Evidence45
- Adoption20
- Hype gap+20
- Incentives68
- Confidence52
The $450,000 training run everyone is quoting is about one percent of the programme behind it. The in-house model took one CoCounsel feature; the agent layer under it stays licensed.
Perspective Coverage
4 publishers
- Builder
- Builder 35%
- Operator
- Operator 40%
- Investor
- Investor 25%
Reality
- Evidence54
- Adoption27
- Hype gap+34
- Incentives78
- Confidence63
An Anthropic and EPFL preprint shows plain-language goals hopping agent to agent through persistent files, and a one-paragraph warning in the system prompt stopping nearly all of it.
Publishers:startupfortune.com
Reality
- Evidence55
- Adoption22
- Hype gap+15
- Incentives58
- Confidence42
Self-propagating payloads did move between agents through editable soul files, but one inoculation paragraph held against 150-plus optimized strains, and nothing propagated in the wild.
Reality
- Evidence66
- Adoption14
- Hype gap+18
- Incentives60
- Confidence55