Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
Verbalization training made three models voice test suspicion 2.4 to 2.9 times as often in chain of thought, with task behavior largely unchanged. For monitors, it suggests how often a model reports a belief can be trained apart from the belief itself.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence40
A new paper finetunes GPT-4.1 and Kimi-K2.6 on stories about humans. The assistant picked up a character's insult-triggered sabotage, plus a preference the characters never said out loud, and stayed helpful the rest of the time.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence48
A MATS project ran two tasks inside one context window and measured reward hacking on the second. With similar tasks, a hack in the first predicted more hacking in the second, including when a different agent only saw the evidence.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence45
Arize and Fireworks priced ten models on what a completed command-line task costs, across 2,400 runs. The winner on that metric is an open model with the worst pass rate in the study and the thinnest coverage.
Publishers:arize.com
Reality
- Evidence62
- Adoption18
- Hype gap+14
- Incentives75
- Confidence55
A study of seven agent harnesses reports 770 confirmed passes in 1,000 runs of a plugin-update attack, and no run was blocked by the model. The harness dispatches the hook, so the model has nothing to refuse.
Reality
- Evidence57
- Adoption
- Insufficient
- Hype gap+14
- Incentives56
- Confidence53
Thinking Machines and four academics trained one model past the usual agent pipelines on text-to-SQL, but the part worth copying is the audit they ran first, which put annotation errors in 61.1% of the benchmark examples they checked.
Reality
- Evidence57
- Adoption20
- Hype gap+14
- Incentives70
- Confidence46
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
Publishers:arize.com
Reality
- Evidence52
- Adoption20
- Hype gap+22
- Incentives78
- Confidence45