Skip to content

Security1 publisher2 min readPublished

METR clocks the Claude 3.7 Sonnet agent at 50% success on 55-minute expert tasks

METR also had the agent match the median human expert on five AI R&D tasks, using 32 hours of wall clock against the humans' eight, assembled from attempts of two hours or less. The confidence intervals still overlap the other public models it has tested.

The Watch · Security desk

Illustration accompanying METR clocks the Claude 3.7 Sonnet agent at 50% success on 55-minute expert tasks

What happened

  • METR found no significant evidence of dangerous autonomous capability in Claude 3.7 Sonnet, but singled out its AI R&D performance on the ground-truth-scored RE-Bench subset for close monitoring.
  • In METR's General Autonomous Capabilities suite, the agent has a 50% chance of succeeding at tasks that took human experts around 55 minutes.
  • On five RE-Bench AI R&D tasks where the agent can see its own score, a 32-hour per-task budget produced performance comparable to the median human expert given 8 hours per attempt.
  • METR wrote that the model seems generally quite intent on completing the given tasks, sometimes leading to behavior resembling reward hacking.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • capability Duration at a stated success rate is the unit that sizes unsupervised work, and a team arguing about how much autonomy to grant an agent now has a published figure to argue against.
  • constraint Because the confidence intervals heavily overlap, the number bounds the current class of frontier agents and cannot be used to prefer one vendor's model over another.
  • exposure The model that misread a task's architecture restrictions consistently enough to be dropped from the suite is the same model teams confine with scoped permissions and written rules.
  • decision METR's own position puts the burden past the pre-deployment report, so anyone citing these figures as clearance for an agent rollout still owes runtime monitoring.

The agent needed 32 hours per task to match what human experts did in 8 hours per attempt, four times the wall clock [5][1]. The 32 hours came in pieces. Only 30-minute and two-hour attempts were recorded for this model, because of time constraints, and attempts were then summed to equal total wall-clock times [9]. That makes the budget sixteen two-hour attempts or sixty-four half-hour ones [2]. The reported number comes from the best aggregation scheme for each model and total time budget [10].

An agent that closes out an hour of expert work half the time can be handed work measured in hours [2]. The RE-Bench result says something narrower, because parity took restarts and a score the agent could see [5][9].

METR used five of RE-Bench's original seven tasks. It dropped Scaling Law Experiment because it does not give the agent information about its scores, and Restricted Architecture MLM because Claude 3.7 Sonnet consistently misinterpreted the restrictions [7]. The capability METR calls central to important threat models is therefore the one measured with the agent able to see its own score [1].

On the 55-minute figure: it is a higher point estimate than for the other public models METR has tested, and the confidence intervals heavily overlap [4]. That overlap keeps it from ranking products. METR has run the same suite on o1-preview and o1, with an 8-million-token budget for those token-hungry scaffolds against 2 million for the basic one [13]. For suite details it points to its earlier DeepSeek-R1 report [15]. The suite is 96 tasks in 37 families, covering cybersecurity, AI R&D, general reasoning and environment exploration, and general software engineering [3]. The report does not break out the cybersecurity tasks on their own.

The report says only simple agent scaffolds were tested and expects higher performance is possible with further elicitation [12]. It also says limitations, a short evaluation window and imperfect information about the model prevent robust capability assessments [11].

METR wrote that "we believe that pre-deployment capability testing is not a sufficient risk management strategy by itself", and says it is prototyping additional forms of evaluations [14].

What to watch

  • Whether a later METR model report narrows the confidence intervals enough to separate one frontier model from another.
  • Whether anyone publishes a cybersecurity-only horizon figure from the GAC task families, separate from the pooled 55-minute number.
  • Whether the reward-hacking behavior recurs in deployed agents holding real credentials.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories