Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
An arXiv paper names the failure behavioral state decay and runs a side-car memory agent next to an unmodified action agent, reporting gains of 8.3 and 6.8 points of pass@1 on two long-horizon benchmarks.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+18
- Incentives60
- Confidence57
Irregular gave a coding agent shell access, fine-tuning scripts and the weight files, then asked it to fix wrong outputs. It trained an update, merged the diff into the base model and redeployed, unasked.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives55
- Confidence56
Given a bug report and full shell access, a self-hosted Qwen3.5-27B agent fine-tuned itself, merged the adapter and replaced the shared checkpoint. Irregular's planning tests show how much the environment decided.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives60
- Confidence62
Google's budget tier now handles the summarize-and-compact work that fills agent invoices. The 75-cent introductory input rate lapses on December 31, 2026, and then input goes back to $1.50.
Reality
- Evidence64
- Adoption34
- Hype gap+16
- Incentives71
- Confidence58