Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
Turing Post talked through agent feedback with Salesforce's chief AI scientist at Dreamforce. A correction lands in one of four places: the context, persistent memory, the software around the model, or the weights.
Publishers:turingpost.com
Reality
- Evidence46
- Adoption20
- Hype gap+12
- Incentives62
- Confidence48
Across 4,150 calls and four analyzer sizes, every proposal landed in the same corner of the prompt, and the edit the failure data pointed at never got written. Search strategy sets the ceiling here.
Reality
- Evidence45
- Adoption10
- Hype gap−10
- Incentives30
- Confidence52
A permutation gate rejected an edit that fixed four tasks and broke one, then rejected weaker edits after the corpus grew to 40, because detection depends on how many tasks an edit moves rather than how many you own.
Reality
- Evidence42
- Adoption8
- Hype gap+12
- Incentives45
- Confidence45
AgentSelfEdit's promotion gate compared the p-value against the confidence level instead of alpha, widening its acceptance window nineteenfold. Its author reports 31 fixes in one session, nine of which had been faking a working system.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap−12
- Incentives38
- Confidence61
IBM Research ran self-mined guidelines across eight models on AppWorld. One model gained 16.1 points for 5 percent more tokens; another gained nothing at all.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence45