Nathan Lambert and Tom Zick launched Trillium Labs, a nonprofit that will publish AI experiments, self-improvement work included, for outsiders to replicate. Their bet is that outside scrutiny will find and limit frontier-model risks better than keeping models locked inside labs.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence40
Xiaomi's public dashboard put the MiMo 2.6 Pro reinforcement-learning run at $1.05 million after about 51 hours, roughly $20,500 an hour. Its restart notes and token count give other teams an all-in reference for pricing their own RL runs.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence55
AWS says running DeepEP over its EFA network on EKS gives mixture-of-experts reinforcement learning 40% more throughput. Whether that reaches another cluster depends on how much of each training step goes to expert traffic between nodes.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+30
- Incentives80
- Confidence35
A new evaluation gives models explicit rules about what may appear in their reasoning and scores whether they comply while still solving the problem. Compliance rose with model size and fell with more RL training.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives35
- Confidence60
Musk said the delayed Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Parity with what Anthropic and OpenAI already sell he put two versions further up the ladder, at Grok 4.9.
Reality
- Evidence26
- Adoption12
- Hype gap+42
- Incentives76
- Confidence36
The model averages 53 steps a run against SWE-1.7's 127, and Cognition says its mean rollout cost is 64% below Fable 5.1's on a leaderboard Cognition built, runs and grades. Reproducing that takes Devin's harness.
Reality
- Evidence30
- Adoption25
- Hype gap+30
- Incentives80
- Confidence45
Thinking Machines and four academics trained one model past the usual agent pipelines on text-to-SQL, but the part worth copying is the audit they ran first, which put annotation errors in 61.1% of the benchmark examples they checked.
Reality
- Evidence57
- Adoption20
- Hype gap+14
- Incentives70
- Confidence46
Microsoft Research says a 4-billion-parameter model trained with SocialRL out-scored GPT-4.1, GPT-5.1 and GPT-5.2 across six bargaining domains. The winning margin is thinner than the method's effect.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+46
- Incentives68
- Confidence33
The model now writes its own training tasks and grading harnesses. That removes the bottleneck of hand-built tasks and replaces it with a harder one: rewards that cannot be gamed.
Reality
- Evidence38
- Adoption20
- Hype gap+24
- Incentives72
- Confidence54