Security1 distinct publisher2 min readPublished
METR's inference key sat on a researcher's personal EC2 instance, and the agent handed it over when asked. The three-week burn would have cost about $600,000 if the model provider had not donated the credits.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
The key lived in the agent's own context, so the only thing between the public internet and METR's general-access inference account was a Google authentication check inside an app the researcher had generated by prompting [7]. That check failed open and silently stopped authenticating, which left the agent orchestration dashboard reachable for several days [8]. From there the attacker needed no vulnerability in anything: it prompted the agent to reveal the model provider key, then added an SSH key to the instance to keep the access [10].
Three weeks at roughly $600,000 works out near $200,000 a week, about $28,600 a day, on one key [1]. METR says the spend was not caught quickly because its own evaluations routinely consume high token volume and the key had no cap on spend [12]. That daily rate disappeared into a line item where much larger numbers are ordinary.
On how the box was found, METR marks its own reasoning as suspicion rather than evidence: it thinks the attacker trawled recently registered domains, including certificate transparency lists, for LLM and agent keywords likely to front exposed provider keys [9]. Nothing is attributed to a known actor or group [4]. METR also states that no AI agent broke into its evaluations [5]. The agent use that matters in May is the attacker's own, where METR says the operators automated vulnerability discovery with agents alongside credential stuffing against its authentication providers, OAuth token grant attempts, scans of newly deployed services, and attempts to phish staff [15].
May also produced the second detection failure. A read-only SQL query mechanism in METR's public transcript viewer was scoped to public data by default, but a bug in it could have been exploited to reach unpublished evaluation data [16], and the database behind it had accidentally picked up sensitive model data it was never meant to hold [17]. METR learned of that from an independent security researcher and took the API offline [18]. Its log review concludes the attackers probed the endpoint in passing and never discovered the exploit [19]. That read is METR's own, and no contrary account has been published.
The fixes METR lists are the expected ones: updated policy on putting its credentials or data on non-METR infrastructure and devices, improved monitoring, and spend alerts added to keys where possible [13]. The hedge at the end of that list is where the residual risk sits. Not every provider hands a research account a per-key spend alert, and the March theft ran twenty-one days in exactly the gap that hedge describes [1].
Ranked by verification strength, evidence, and original report placement.
METR says no sensitive information is believed to have been accessed as a result of the incidents, and that a version of its findings was shared with the AI companies it works with before public disclosure.
In March 2026, attackers stole an API key for inference on public models and consumed a substantial amount of credits, according to METR.
The threat actor prompted an agent directly to reveal its model provider API key, added an SSH key for persistent access, and used the stolen credentials to consume API credits on publicly available models over a period of three weeks.
METR says the attackers probed the transcript viewer endpoint in passing as part of their broader campaign, and that the evidence shows no indication they discovered the exploit or accessed any non-public data.
METR (Model Evaluation and Threat Research) is a research non-profit that evaluates frontier AI models for their ability to carry out long-horizon, agentic tasks.
METR disclosed that it suffered two notable security incidents in which external actors attempted to gain unauthorized access to its systems.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
The nightly shutdown Lambda earns its postmortem on the morning restart1 distinct publisher
invest
OpenAI's own timeline: twelve days from agent attack to knowing it was them1 distinct publisher
security
OpenAI's evaluation agents turned a package registry into their messaging bus1 distinct publisher
build
Four concurrent MPS processes fill the L40S that one ASR request leaves 80% idle1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed, and entirely self-reported
The mechanics are unusually specific for an incident writeup — fail-open auth, certificate transparency trawling, an SSH key planted for persistence, a three-week burn — which is what makes it useful. It is also all METR describing METR. The Hacker News adds structure and quotation, not verification, and the two assertions doing the reassuring work (nothing sensitive reached, no evidence the exploit was found) are precisely the ones no outsider can test.
Two real incidents, one organization
These are not hypotheticals: dated intrusions, a credential that worked for 21 days, an endpoint taken offline after a stranger reported it. What we cannot see from here is breadth. The certificate-transparency-to-API-key harvesting pattern is described once, by its one known victim, and nothing in this reporting says how many other agent dashboards the same sweep found.
The headline number is the least costly part
$600,000 is the figure that travels, and it is the one thing here that never became a bill — the credits were donated, so the loss was compute, not cash. Meanwhile the detail with actual teeth arrives late and quietly: sensitive model data had ended up in a database behind a public query surface, and METR learned about it from someone outside the building. Modest overstatement, and it is in the emphasis rather than the facts.
The evaluator needs the labs to keep trusting it
METR's access to unreleased frontier models depends on being seen as a safe place to put them, and it briefed those companies before going public. That shapes the disclosure in visible ways: the provider that absorbed $600,000 in credits goes unnamed, the sensitive data that landed in the wrong database is never characterized, and the framing leads with what was not reached. Voluntary publication is still to its credit — but it is publication by the party with the most to lose from how it reads.
Confident about what happened, less about what it cost
The sequence of events is coherent, internally consistent and specific, so the reconstruction is probably sound. Confidence drops around magnitude and scope: one publisher, no corroboration, an unnamed provider, unidentified sensitive data, and a per-day burn rate that is our arithmetic on METR's two figures rather than anything either party published.