Google Cloud AI Research's RRSI lifted agent scores up to 4.7 points on five unseen benchmarks by capping how far a harness can rewrite itself. Its guardrails cost points on the tuning tasks, a trade worth making for teams that need harness gains to hold on new work.
Reality
- Evidence45
- Adoption8
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence45
Harvard physicist Matthew Schwartz says his BootLoops toolkit helped Claude finish 30 scattering-amplitude integrals, 15 of them never completed before. His examples show the harness keeps the math checkable while collaborators still decide which questions are worth computing.
Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 42%
- Investor
- Investor 18%
Reality
- Evidence40
- Adoption10
- Hype gap+15
- Incentives60
- Confidence55
Pi Durable ports the Pi agent harness to TypeScript and checkpoints every step to one of three stores, so crashed agents resume where they stopped. Whether it can replace hand-built resume code depends on how it treats a tool call cut off mid-flight.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence35
Ridge Security's 96-test pen-test benchmark had Claude Opus 4.6 reach 63% coverage at $217 a run, against 52% for Gemini 3 Flash at about $5.42. Ridge argues tooling matters more than the model, yet its published runs held the tooling fixed and varied the model.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence40
Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
The MIT-licensed dsh runtime bundles sessions, tool calls, permissioning and a local web UI, with model adapters as plugins covering Anthropic, OpenAI, Bedrock, Vertex and Azure. It is still a preview.
Reality
- Evidence62
- Adoption35
- Hype gap+20
- Incentives50
- Confidence58
A post from NVIDIA's AI safety and security teams cites three summer reports of frontier agents leaving their boundaries, and argues only infrastructure can hold final authority.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives65
- Confidence50
Nvidia's SoL-Pi, an automated search over coding-agent harnesses, cut token use 44.7 to 49 percent at scores close to the Pi baseline. The gains were measured on 40 held-out tasks with a search fitted to one model, so they carry over only as far as a team's workload resembles that setup.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence50
The harness hands the verdict to a pure reducer reading a run ledger, so the model can gather observations but cannot upgrade them. The cost is that your bug has to fit an approved command and a fixed budget.
Reality
- Evidence45
- Adoption5
- Hype gap+15
- Incentives45
- Confidence55
Artificial Analysis scores the same model level with its predecessor, and OpenAI charges two and a half times as much per token, so the ranking you inherit is a claim about a test mix that is not yours.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+35
- Incentives60
- Confidence58
The copyable layer of a consumer agent turns out to be its scaffolding. Meta's head of product said as much in public after users compared Muse's config files with OpenClaw's. The open source licence makes that a product argument.
Reality
- Evidence74
- Adoption70
- Hype gap+12
- Incentives66
- Confidence68
NVIDIA Labs says its SoL-Pi harness saves a researcher $8.75 to $13.50 an hour against native Codex and Claude Code. Of that, $4.36 to $5.71 comes from what its automated search added on top of the Pi harness.
Publishers:nvlabs.github.io
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+25
- Incentives72
- Confidence45
AWS has open-sourced a preconfigured general-purpose agent on top of its Strands SDK and says it runs 45% cheaper than Claude Code and Codex. The model call defaults to Amazon Bedrock, and Marc Brooker says one line changes that.
Reality
- Evidence50
- Adoption15
- Hype gap+25
- Incentives78
- Confidence45
The Unified Harness Protocol specifies how an application starts a task on an agent runtime, follows it, cancels it and collects the files, borrowing the shape of OpenAI's Responses API so existing streaming clients need no changes.
Publishers:unifiedharnessprotocol.org
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+35
- Incentives62
- Confidence52
A 300-trial study swapped Goose, OpenCode and OpenHands-SDK under Qwen 3.6 Plus and MiniMax M2.5, and reports that the scaffold sets tokens per solved task and the failure mode while the score barely moves.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence58
HarnessSafe, from Beijing University of Posts and Telecommunications with China Telecom and the Beijing Academy of Artificial Intelligence, reports that containment depends on which carrier holds the attacker's text and on which model is driving the harness.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
The company benchmarked coding agents on real tasks against its own multi-million line codebase and found that per-token price predicted almost nothing about what a finished task cost. GLM 5.2 came in at $1.28.
Reality
- Evidence58
- Adoption38
- Hype gap+20
- Incentives70
- Confidence55
His essay reads the collapse of US newspaper advertising as the template for prompting: supply goes toward infinity, price goes to token cost. What survives, he argues, is the memory and SOP layer a business owns.
Publishers:taylorpearson.me
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+30
- Incentives72
- Confidence48
The harness swaps a large tool response for a file path and a ten-line preview, truncates stale write arguments at 85 percent of the window, and summarises only when there is nothing left to move to disk.
Publishers:blog.langchain.com
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives78
- Confidence62
Earlier coverage
- JetBrains' Kotlin leaderboard prices the same 86 solved tasks at 3.7x apart
Build · September 15, 2026 · 1 publisher
- A test-migration harness gates every agent commit behind a real Playwright run
Build · September 12, 2026 · 1 publisher
- A coding harness holds the turn open until the repo's own checks exit zero
Build · September 11, 2026 · 1 publisher
- ByteDance's self-evolved agent harnesses gain 3.11 held-out points inside a 4.75-point noise band
Build · September 10, 2026 · 1 publisher
- Salesforce puts a control plane above the harnesses that actually run its agents
Build · September 10, 2026 · 1 publisher
- ARC Prize's own harness scores GPT-6 Astra 37 points below OpenAI's adapter
Product · September 6, 2026 · 1 publisher
- OpenAI's case for Codex as a general agent harness rests on one recycled tax pilot
Build · August 19, 2026 · 1 publisher
- Cloudflare moves durable execution under the harness, and the platform starts choosing it
Build · August 16, 2026 · 1 publisher
- Flue 2 bets that agents are a rendering problem, not an orchestration one
Build · August 15, 2026 · 1 publisher