JetBrains and Lund University propose review tools that flag per-segment risk in AI-written code, a design shaped with 17 practitioners and a 43-person survey. They argue diff review cannot keep up with agent-sized changes unless tools show reviewers where to read closely.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
Researchers found an auto-displayed AI answer cut 'I don't know' responses from 35 percent to 1 percent, though the model was mostly wrong. A 10-cent penalty for wrong answers still left abstention at 7 percent, so review tools need more than an abstain button.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence50
Three Northeastern researchers describe chatbot vocabulary and sentence habits moving into human prose. The measurement behind that description is a preprint on how people and ChatGPT copy each other mid-conversation.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence45
Manchester and Durham researchers had 390 people rate comfort messages without knowing who wrote them, and the AI text scored higher for anger and fear. Human messages carrying the same practical advice were judged just as supportive.
Reality
- Evidence45
- Adoption30
- Hype gap+12
- Incentives55
- Confidence50
Applicants who believed software would score their one-way videos stretched the truth more often, the rating agent scored them no lower for it, and telling them what it measured brought the exaggeration back down.
Reality
- Evidence54
- Adoption
- Insufficient
- Hype gap+18
- Incentives58
- Confidence47
The MIT and Motional method, published in Nature, routes a pretrained planner's final decision through concepts a person can read, which is how the team claims faithful explanations and unchanged driving at once.
Reality
- Evidence62
- Adoption20
- Hype gap+12
- Incentives62
- Confidence68
In a 1,155-person study, one participant in 600 found the optimal play unaided. With an AI demonstrator seeded in the first generation, the strategy survived in nine of 15 groups.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+24
- Incentives55
- Confidence54