Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
Intel has added NVIDIA's Apache-2.0 OpenShell policy layer to its enterprise agent toolkit, putting sandboxing and kernel-level agent controls on Xeon. The feature ships switched off, so what it is worth to Intel depends on how many agent teams turn it on.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives60
- Confidence45
Google open-sourced AX, an Apache 2.0 runtime on Kubernetes that checkpoints idle agents and is designed to resume them in under a second. Its savings depend on how long each agent waits on models, tools or people.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
AWS has open-sourced a preconfigured general-purpose agent on top of its Strands SDK and says it runs 45% cheaper than Claude Code and Codex. The model call defaults to Amazon Bedrock, and Marc Brooker says one line changes that.
Reality
- Evidence50
- Adoption15
- Hype gap+25
- Incentives78
- Confidence45
Anthropic's open-source audit framework now runs a classifier over every auditor turn and rewrites anything a real deployment would not produce. The tuning targeted models that say out loud they are being tested.
Publishers:alignment.anthropic.com
Reality
- Evidence55
- Adoption35
- Hype gap−10
- Incentives75
- Confidence45
Ao Qu and collaborators open-sourced Reef on September 15th, an OpenAI-format inference server that stamps each response with a record ID, matches later feedback to that ID, and publishes a retrained artifact only if its evaluation stage accepts it.
Reality
- Evidence48
- Adoption10
- Hype gap+12
- Incentives32
- Confidence60
CauterRule's own field test puts the undecided bucket above pass and fail combined, and the report names the matcher that produced it as its top calibration target. The per-model rates end up measuring trigger phrasing.
Reality
- Evidence32
- Adoption12
- Hype gap+10
- Incentives82
- Confidence30
Nvidia's Apache 2.0 beta seizes the port Ollama or LM Studio was using and forwards each call to whichever machine is free, which makes local throughput a function of how many PCs you own. Hands-on testing found it serving an engine Nvidia does not list.
Reality
- Evidence55
- Adoption15
- Hype gap+10
- Incentives62
- Confidence48
Headlong is Apache-2.0, so a team can read the loop before running it. What nobody can read is a benchmark, and the shared thought stream is also the privacy problem.
Reality
- Evidence34
- Adoption16
- Hype gap+22
- Incentives66
- Confidence42
DeepSeek Harness is an MIT-licensed developer preview in which the model adapter, tool registry, session log and agent loop are all plugins. The interesting part is what that implies for evaluation.
Reality
- Evidence42
- Adoption12
- Hype gap+28
- Incentives55
- Confidence38
The best of 12 models passed 65.36% of 507 business tasks on the first attempt and only 25.25% across all 20 trials. Agent procurement should be priced on the second number.
Reality
- Evidence40
- Adoption16
- Hype gap+12
- Incentives62
- Confidence38
An open-source stack pairs a deterministic Minecraft reimplementation with seed-level provenance, so a reinforcement learning result can be replayed instead of reconstructed.
Reality
- Evidence42
- Adoption10
- Hype gap+12
- Incentives34
- Confidence45