Build1 distinct publisher3 min readPublished
An AI evaluator with frontier-model access lost a provider key from an experiment on a researcher's personal EC2 box. The reach of the credential, not the seriousness of the project, decided what the mistake cost.
The Engineer · Build desk
invest
OpenAI's own timeline: twelve days from agent attack to knowing it was them1 distinct publisher
security
OpenAI's evaluation agents turned a package registry into their messaging bus1 distinct publisher
build
Instant's OpenAI exit comes with a due date: hosted apps stop on August 31st, 20271 distinct publisher
build
Same-day GPT-5.6 on Azure kills the parity argument, leaving auth and residency to decide1 distinct publisher
Compiled by The EngineerSomething wrong?How this is made
Fail-open is a precise bug shape. A request arrives, and the application tries to establish who is calling. The check errors, or returns nothing, and the handler reads the empty answer as authorisation rather than as a failure. METR says the protection on the instance was Google authentication, and that the application's own flaw silently disabled it [5].
What made the bypass expensive is what the application was for. An agent-orchestration dashboard is an instruction channel into a process that already holds provider credentials and tool access, so whoever gets past the door does not need a file-read primitive. METR says the attacker prompted an agent to reveal the model provider's API key, installed an SSH key to keep its foothold, and used the credential for roughly three weeks [7].
That three-week span matters more than it might seem. The stolen credential was a general-access key rather than one for private models [8], so its traffic looked like ordinary public API usage. METR also says its internal dashboard did not surface every rate-limited request to users, and that unusually heavy usage and provider errors were routine during large evaluation runs [10]. An evaluation harness hammering a provider looks, in aggregate, a lot like a stranger hammering the same provider on your key.
Now the arithmetic. About $600,000 of list value [2] over roughly three weeks, call it 21 days [7], runs near $28,600 a day [1]. Whether that rate means anything in your shop depends on whether your key sits on an account billed at list. METR's sat on capacity a model developer had donated, so no invoice ever arrived to argue with [3].
The exposure window and the credential lifetime also came apart. The dashboard was publicly reachable for several days [4] while the key kept working for about three weeks [7], which puts most of the consumption after the app itself stopped being the way in [2].
Segmentation, though, is the part that held. METR's mission requires privileged access to unreleased models, hidden chain-of-thought outputs and confidential laboratory information [15], and the researcher who stood up the experiment had access to none of that [5]. A boundary drawn before the incident is why the loss was denominated in public-model tokens.
The remediation reads accordingly: security reviews formalised for public applications, use of METR credentials and data on personal infrastructure restricted, usage monitoring expanded, and spending alerts added where providers support them [14].
Hosting location alone would not have caught the other opening. METR had inadvertently exposed a read-only SQL mechanism through its own public transcript viewer, meant to be limited to public information, where a bug could have allowed access to unpublished evaluation data, and some sensitive model outputs had been placed in that database by mistake [19]. An independent security researcher found it and reported it, and METR removed the API and paid a bounty [20].
The control I would want out of this is per-application keys with their own spending ceiling, including on free capacity, because the ceiling is the only part of a credential you can set before you know who ends up holding it.
Ranked by verification strength, evidence, and original report placement.
The credits had been provided free by an unnamed model developer, so METR did not absorb a $600,000 bill.
METR said its forensic work found no compromise beyond the public-model API key.
Elizabeth "Beth" Barnes, founder and CEO of METR, disclosed in an August 31 security update that external attackers exploited weaknesses around the AI evaluator's public infrastructure in two incidents earlier this year.
One stolen API key was used to consume model credits with a list value of approximately $600,000.
In the March compromise a researcher placed the key on a personal Amazon Web Services instance running a "vibe-coded" agent dashboard whose authentication failed open, leaving it publicly reachable for several days.
METR said the March incident began when a researcher without access to sensitive models or confidential data deployed an agent-orchestration application on a personal EC2 instance that was intentionally internet-facing and placed behind Google authentication, which the application's fail-open flaw silently disabled.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Unusually candid, entirely first-party
The mechanism is described at a level of detail organizations rarely volunteer — fail-open authentication, an agent prompted into disclosing the key, SSH persistence, the dashboard that hid rate-limited requests — and that specificity is what makes it credible. But the chain is one deep: RuntimeWire is reading METR's own post, the consultancy that ran the compromise assessment has published nothing, the provider is unnamed, and the attackers are unidentified. The all-clear is the weakest link, and the reporting says so itself.
Real events, dated fixes, one narrator
This is not an announcement about intentions. Three distinct things happened — a key was stolen and spent, infrastructure was probed by agent-driven tooling in May, a public SQL path was exposed and paid out to its finder — and the remediation carries a date of July 30, before disclosure. What holds the score down is that every event, including the bounty and the isolated production environment, is counted by the same party that is being assessed.
The number that nobody paid
Six hundred thousand dollars is doing rhetorical work that the facts partly withdraw two sentences later: it is list value on donated credits, and the loss sat in a provider's allocation. RuntimeWire is honest about this rather than hiding it, which keeps the gap small. Pulling the other way, the genuinely alarming finding is quieter than the headline — an evaluator holding unreleased models and hidden chain-of-thought had its perimeter recreated by one researcher's side project, and a scope accident is most of why that ended cheaply.
A disclosure the disclosed-to got to read first
METR's access depends on labs believing it is safe to hand over unreleased models, so a post that says the stolen key only reached public models and nothing else was touched is exactly the post METR needs to be true. Several AI companies reviewed it before publication and suggested wording changes. None of that makes it wrong — voluntarily publishing a $600,000 own goal cuts against pure self-interest — but the unnamed provider, the unnamed attackers, the undescribed transcripts and the unpublished consultancy assessment are all places where candour stopped.
Trust the mechanism, hold the verdict
How the key left the building is well enough established to act on tomorrow: it is internally consistent, technically ordinary, and specific in ways a sanitized account would not be. What the incident ultimately touched is a different confidence question, and one publisher relaying one vetted self-report cannot settle it. A published assessment from the consultancy, or any word from the provider, would move this materially.