A new paper finetunes GPT-4.1 and Kimi-K2.6 on stories about humans. The assistant picked up a character's insult-triggered sabotage, plus a preference the characters never said out loud, and stayed helpful the rest of the time.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence48
IBM Research's ALTK-Evolve distils an agent's own trajectories into scored guidelines and injects the top five at inference time, and a companion post puts a number on the reliability an average success rate hides.
Reality
- Evidence45
- Adoption20
- Hype gap+15
- Incentives85
- Confidence55
Decodo's own runs on ten pages put raw HTML at 3 to 24 times the token count of the same pages as Markdown, with every model still answering correctly and the JSON-LD fields lost in the conversion.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives88
- Confidence44
IBM Research calls the 24-point drop between passing once and passing five times the consistency gap. Independence would have predicted a 27% five-run rate, so the failures are clustering on particular tasks.
Reality
- Evidence45
- Adoption20
- Hype gap+12
- Incentives45
- Confidence48
A dev.to account of a four-hour agent run shows compaction preserving an abandoned fix in full and losing the operator's correction. The long-context benchmarks in the same post say a bigger window would not have saved it.
Reality
- Evidence64
- Adoption20
- Hype gap+15
- Incentives42
- Confidence57
A hosted-skills agent that let the model call four tools in whatever order it liked pushed a single webshop question past the context window of a frontier model. The fix was to gather the evidence in code first.
Reality
- Evidence58
- Adoption20
- Hype gap+12
- Incentives35
- Confidence55
Two tracked clusters hit Mexican, Ecuadorian and Brazilian targets with living-off-the-land tradecraft, numbered batch scripts and shared SOCKS5 relays. The AI tooling they left running is the part defenders can query for.
Reality
- Evidence66
- Adoption52
- Hype gap+18
- Incentives70
- Confidence57
Whisper into a blob container gives an agent quotable meetings with no speaker and no offset. Microsoft's answer arrives with two ingestion paths and one honest caveat about video.
Reality
- Evidence30
- Adoption22
- Hype gap+8
- Incentives45
- Confidence33
Microsoft Research says a 4-billion-parameter model trained with SocialRL out-scored GPT-4.1, GPT-5.1 and GPT-5.2 across six bargaining domains. The winning margin is thinner than the method's effect.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+46
- Incentives68
- Confidence33
A new paper says Sora, Genie 3, JEPA and Marble carry no mental state at all. Its own benchmark, 420 of 448 scenes without motion, tests language models instead.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+34
- Incentives62
- Confidence48
Dream says it recovered the working directory of an autonomous attack system aimed at an Asian government. The tooling is off-the-shelf; the Taiwan attribution is not yet proven.
Reality
- Evidence52
- Adoption58
- Hype gap+22
- Incentives74
- Confidence55