Security1 publisher2 min readPublished
METR let Anthropic review and edit its Claude Opus 5.5 evaluation summary before sign-off
METR's predeployment report on Claude Opus 5.5 names its five test tasks and the ten business days of API access behind them. It also states that the work was not meant to verify Anthropic's policy thresholds and did not assess alignment.
The Watch · Security desk
What happened
- METR's preliminary evaluation asked whether Claude Opus 5.5 would dramatically accelerate AI R&D at Anthropic, and whether AI had already dramatically accelerated the model's own development.
- METR concluded the model is a modest improvement over Fable 5.1, and is unlikely to be able to fully automate AI R&D.
- One input was a separate METR team's report on acceleration inside Anthropic, which shared its conclusions with the authors but not the evidence or reasoning behind them.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- constraint The report cannot be used to show that a model cleared a policy threshold or holds an alignment property, so it will not support a risk register that cites it for either.
- decision Anyone setting AI risk posture from third-party evals has to judge what two calendar weeks of API access across five tasks can actually establish.
- precedent METR put its access terms and the vendor's review rights in the report itself. The next evaluator will be compared against that.
- contradiction METR's tasks show an incremental gain while Anthropic's questionnaire puts the model on the Mythos-level trend on its own internal measure. Only the first is checkable from outside.
Five tasks in ten business days averages two business days per task, and that window covers setup, runs and write-up [5][6][18]. Two of the five produce output a tester can check directly: Budget NanoGPT Speedrun and Gaming Bot [11]. LMCA and Sunlight are what METR calls harder-to-verify [11]. Train a Program is in the task list but appears in neither group where METR states that Opus 5.5 improved on Fable 5.1 [20]. METR calls the whole exercise a preliminary evaluation [21].
The independence note is METR's own. "We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text," the report said [3]. METR signed off on the final wording, and that wording is the text carried in the Claude Opus 5.5 system card [2]. The work was done under an unpaid agreement for AI R&D assessment [1]. The note does not say whether Anthropic asked for changes [22].
Two exclusions carry more weight than the conclusions for anyone citing this in a risk register. The report states the work "was not meant to verify claims about compliance with any specific threshold from Anthropic's policies" [7], and that it "does not attempt to assess whether Claude Opus 5.5 has or does not have particular alignment properties" [8].
Among the inputs is a report METR cannot show: a separate METR team, working with elevated access, assessed AI R&D acceleration inside Anthropic and passed its conclusions to the evaluation team without the supporting evidence or the details of its reasoning [15]. METR uses those conclusions as an input and does not argue in defense of them [15]. The remaining inputs are Anthropic's answers to a capability questionnaire, an interview with an Anthropic researcher, and METR's own Frontier Risk Report [17].
On its five tasks, METR is reasonably confident that Opus 5.5 is a modest improvement over Fable 5.1 and not a huge leap [10]. The report describes the gain on its quantitative evaluations as incremental [12]. The model still shows qualitative weaknesses an expert human is unlikely to exhibit on hard, long-horizon tasks or open-ended reasoning [13]. METR expects acceleration slightly higher than from Fable 5.1 and says the model is unlikely to be able to fully automate AI R&D, an expectation it calls highly uncertain [9][23]. The claim that Opus 5.5 continues the Mythos-level trend on Anthropic ECI came from the questionnaire and the researcher interview, not from the five tasks [14].
What to watch
- The further public outputs METR expects from the separate elevated-access investigation of AI R&D acceleration inside Anthropic.
- Whether a later METR summary lists the edits a vendor requested, or publishes the pre-review draft alongside the signed-off text.
- Whether Anthropic ECI numbers are published anywhere a reader can check the Mythos-level trend claim outside the questionnaire.