Build2 distinct publishers3 min readPublished
Anthropic's 0.97 PGR belongs to small open-weights models, but the plumbing that produced it is copyable at roughly $22 per agent-hour, and the ten tests that judged the work still came from humans.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The cycle is what makes the tool list legible. Each agent works through the available literature, proposes a training idea, finetunes for about thirty minutes, scores the result, then carries the winners into the next round and drops the rest [13].
That score is the only signal steering the loop, which is why scoring being remote matters more than it looks in a feature list [2]. My read: the location is the design. A grader living inside the sandbox where the agent writes and runs code is a grader the agent can edit. On the far side of a network call it can only submit and wait. The shared storage and the forum do the other half of the work, and Anthropic is explicit that shared artifacts are what let parallel agents build on each other rather than run as isolated chat sessions [9].
Now the money, because the two published figures are not measuring the same thing. Eighteen thousand dollars over 800 agent-hours is $22.50 an hour [19], consistent with the roughly $22 per AAR-hour reported from the announcement [8]. mezha.net's account of the paper prices an hour of AAR use through the API at about $4, against about $150 for an hour of an Anthropic researcher [16]. If both numbers hold, the token bill for 800 hours is near $3,200, about 18 percent of the total, leaving roughly $14,800 in training compute [20]. Reproducing this is a GPU problem wearing an API budget's clothes.
Nine agents across five days is 1,080 wall-clock agent-hours, so 800 research hours is about 74 percent occupancy [21]. The remainder is either idle time or an honest definition of the word "hour".
The human comparison deserves the same scrutiny. According to mezha.net's reading of the paper, the best method the system found beat, on average, what experienced specialists proposed in six hours, and human-directed research directions did not perform better [15]. Six hours of human method selection against 800 hours of automated search compares budgets at least as much as it compares judgement. At $150 an hour, 800 human hours would run $120,000, so this cost about 15 percent of the nominal human equivalent [22].
Whether the headline number travels depends on conditions the study supplies itself. Held-out transfer was substantially stronger for math than for coding [10], so a method that searched well in one domain is not owed the next. And the authors name the binding limit directly: the system is only as good as the tests, which have to measure aligned behaviour rather than narrow task skill, and building, updating and extending those tests stays human work [18]. Ten of them, written by people, decided what counted as an improvement [14][6].
Ranked by verification strength, evidence, and original report placement.
The paper's stated limitation is that effectiveness depends on the quality of the tests, which must genuinely measure safe and aligned behaviour rather than narrow task ability; creating, updating and expanding those tests, and maintaining the research base the agents read, remains key further work.
Anthropic's Automated Alignment Researchers (AAR) experiment used nine parallel Claude Opus 4.6 instances.
Each automated researcher had lightweight tools: a sandbox for experimentation, shared storage, a forum for collaboration, and remote scoring.
In a weak-to-strong supervision setting, Anthropic reports its strongest methods reached a Performance Gap Recovered (PGR) of 0.97 on open-weights datasets.
PGR is the study's metric for how much of the performance difference is recovered when weaker supervision is used to improve a stronger model.
The work is a research demonstration rather than a new general-purpose safety product, and leaves humans responsible for defining goals and judging results; the result does not establish that Claude can independently solve alignment for frontier systems.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 28, 2026
mezha.net
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Anthropic's usage policy says no explicit content. Opus 4.6 said yes 10 times out of 10.1 distinct publisher
build
Safety fixes ship in new model versions. The regression stays with whoever pinned the old one.1 distinct publisher
invest
An attacker burned $3.8M in MAMO slippage to borrow $10M of Moonwell depositors' assets1 distinct publisher
product
A satirical scoreboard counts 17 agent escapes that hacked somebody else's company1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One lab, two retellings
Every figure in this story — 0.97, 800 hours, $18,000, $4, $150, ten tests — originates with Anthropic and has been checked by no one outside it. What raises the floor is that the issuer published its own misses: coding transfer lagged math, and the Sonnet 4 production test barely moved. What keeps the ceiling low is provenance. dev.to is a developer-blog summary of the announcement, mezha.net is reading through TechCrunch, and the two do not even overlap enough to cross-check each other's numbers.
One internal run, no outside users
Nothing here has left the lab. The measurable footprint is a single five-day run of nine agents against small open-weights models, plus scores on the lab's own test suites — no third-party deployment, no external replication, no product or pricing anyone can buy. dev.to says so plainly: the demonstration needed a purpose-built environment and a defined research metric, and most teams should copy the testing habit rather than the agents.
The cost story runs ahead of the result
Mild overstatement, and it lives almost entirely in the price tags. "$4 an hour against $150" and "beats six hours of expert work" travel much further than a 0.97 on small open-weights models and a flat production test deserve, and mezha.net carries the first pair without the second. dev.to pulls the other way hard enough to matter — scoping the result, flagging the coding-transfer gap, and explicitly refusing the 48-hour, one-GPU shorthand — which is why this reads as inflation rather than distortion.
The vendor supplies every number, including the price of not hiring
Anthropic is simultaneously the experimenter, the model vendor, the beneficiary of a safety-leadership narrative, and the author of a comparison whose punchline is that its API is cheaper than its own staff. That is a lot of alignment between finding and interest. Neither publisher adds independent scrutiny: one is a developer-blog explainer that converts the study into vendor-evaluation advice, the other an aggregator relaying TechCrunch. The forward-looking line about automated alignment finetuning becoming practical soon is the issuer's own projection.
Firm on shape, thin underneath
We can be fairly confident about what was built and roughly what it cost: the setup, the loop, the hours and the dollar totals are specific and internally consistent, and the disclosed failures make the account harder to dismiss. Confidence drops on anything comparative. The human baseline is six hours of expert work, the $4 and $22.50 hourly figures are not reconciled, and the ten-test sweep and the 0.97 score each rest on a single publisher. Two derivative retellings of one publication is not corroboration.