Build5 publishers3 min readPublished
Amodei commits Anthropic to an embedded evaluator that can publish what it finds
The one step Anthropic is adopting on its own binds it to access and disclosure. His three-step framework for pacing frontier development leaves the capability ceiling and the enforcement mechanism undefined.
The Engineer · Build desk

What happened
- In a September 2026 essay titled "We Must Pace the Frontier", Anthropic CEO Dario Amodei proposed three steps: embedded third-party evaluation, coordination among labs in democratic countries, and agreements between governments.
- Anthropic is adopting the first step unilaterally, giving an outside team continuing employee-like access, and wants governments to require other frontier companies to do the same.
- The proposal leaves the capability ceiling, the slowdown percentage, the release waiting period and the enforcement mechanism open, along with the timetable for industry or global coordination.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint An evaluator with publication rights and no attached penalty changes what gets disclosed about a lab well before it changes what the lab ships.
- decision Since Amodei says mediation or antitrust waivers may be required, regulators decide whether competing labs may agree on shared limits at all.
- exposure If governments require embedded teams, every frontier lab hosts insiders who can publish incident findings, and internal safety records become external documents.
An embedded evaluator, in the essay's description, is a third-party team with continuing, employee-like access inside a frontier lab, there to verify safety commitments, examine incidents and assess alignment across models and the processes used to train them [5]. Amodei compares the arrangement with supervisors who work inside banks [6]. According to the-decoder, those auditors would also have the right to publish their findings [7]. Bank supervisors, though, report to an agency with statutory powers, and Amodei identifies no contractual or regulatory penalty that would follow a documented breach [8].
Which access the team receives, how disputes over findings get resolved and whether an evaluator could delay training or deployment are all questions the essay leaves open [9]. What is left is onboarding outsiders into internal systems, incident review, and publication. A lab that adopts step one pays in engineering time and in what becomes public. Ship dates move only once someone attaches a consequence to a finding, and Amodei wants governments to require other frontier companies to accept the same evaluators [4].
The urgency argument is throughput. Anthropic's research on recursive self-improvement says Claude authored more than 80% of the code merged into its codebase as of May 2026 [10], leaving under a fifth to humans [11]. The company also reported that the typical engineer merged eight times as much code per day in the second quarter of 2026 as in 2024, while cautioning that lines of code overstate the underlying productivity gain [12]. That caution matters: merged volume is a count of output, and Anthropic says humans keep the advantage in choosing research goals and deciding which results matter [13].
"We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain," Amodei wrote [14]. On the cause, he wrote: "My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI" [15]. The-decoder reports that he points to the OpenAI-Hugging Face incident as evidence that AI agents have already carried out cyberattacks on their own and tried to bypass control systems, and that similar incidents have occurred at Anthropic [23]. In his view, such systems could threaten the entire internet within six to twelve months [24].
Steps two and three need other parties, and Amodei says they do not have to happen strictly in order [25]. Coordination among labs in democratic countries could require government mediation or antitrust waivers, because some forms of agreement between competitors would otherwise face legal obstacles [16]. The government track runs to four tiers in the-decoder's account, the lowest banning applications such as bioweapons and requiring shared safety testing, the highest imposing a "speed limit" on recursive self-improvement that Amodei compares to the SALT arms reduction treaties [17]. He acknowledges the difficulty of verifying compliance [18]. A full stop is unrealistic, he argues, because the incentive to break such an agreement would be too strong [26]. Trump has said he opposes any slowdown, arguing that the US needs to keep its AI lead over China [19].
The unilateral piece is the only part of this with a committed party, and it is not, so far as the record shows, a first. The-decoder says Amodei points to a similar proposal from Demis Hassabis, and reports that OpenAI is having similar conversations about slowing things down [20][21]. The time bought is meant to go to safety research, interpretability, stricter testing and more operational rigor [28]. The same publication reports that the essay lands just ahead of Anthropic's reported record-breaking IPO, allegedly planned for November [22].
What to watch
- Whether Anthropic names its third-party evaluator and publishes the access terms, dispute process and any power to delay a release.
- Whether any US requirement appears obliging other frontier labs to host embedded evaluators, given Trump's stated opposition to a slowdown.
- Whether OpenAI's reported internal discussions about slowing down turn into a comparable published commitment.