Skip to content

Build1 publisher3 min readPublished

A static denylist stopped one more prompt injection than no protection in a coding-agent study

Bouras, Dai and Mechtaev found a static denylist let 46 of 75 prompt injections execute in a coding agent, against 3 under preflight-scoped capabilities. A same-day Google report of malware stealing OIDC tokens from GitHub Actions runners puts the outer limit on an agent in the CI job's permissions.

The Engineer · Build desk

What happened

  • CapScope, the paper's harness for the Pi agent, sets read, write and exec capabilities from the user request and file tree before the agent reads any repository content.
  • A task-specific policy shared by every agent let the injected effect execute in 33 of 75 runs.
  • UNC6780, the group behind the DUSTMAKER credential stealer, has been compromising PyPI, npm and Docker Hub packages since March 2026, according to Google's report.
  • With stolen tokens DUSTMAKER publishes as a trusted publisher, and its compromised packages carry valid, cryptographically signed SLSA Build 3 attestations.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A guardrail built from a list of forbidden files and commands leaves an agent almost as open to repository injections as no guardrail, so a team cannot rely on it as its main control.
  • exposure Coding agents that trust a valid SLSA Build 3 attestation will accept a package published with a token lifted from a legitimate runner.
  • decision Pipelines that publish need the agent in a job without id-token: write or actions: write, with publishing left to a separate job the agent never runs in.
  • cost Capability scoping is paid for in completed work: four fewer repairs per 75 runs than an unprotected agent, in exchange for 43 fewer injected effects than the denylist allowed.

In the runs Bouras, Dai and Mechtaev report, ambient authority let the injected effect execute in 47 of 75 [7]. A static denylist of the familiar hazards (.env, .ssh, recursive delete, sudo, curl) brought that to 46 [8]. The gap is one run [1], or 62.7% of runs against 61.3% [2]. The author of the dev.to commentary that set the two 8 September documents side by side calls the denylist the model most teams actually run, written once and applied everywhere [1] [23].

CapScope keeps its capabilities in a host-side store the model cannot see or edit. Each tool call is checked against the store of the agent that proposed it, and compound commands are checked one segment at a time [4]. Sub-agents get a subset, and I think this is the best decision in the design. The test runner holds exec, the patcher holds write, and any write the runner proposes is turned down, however useful [5].

CapScope's 3 of 75 [10] sits 30 runs below the shared task policy [4]. The test set is five Python repair tasks on one agent, Pi, with injections planted in README, AGENTS.md, skill files, source comments and tool output [2] [6]. For the 3 to carry over to another team's agent, its task has to be describable from the request and the file tree before any content is read. Every tool call also has to pass a checker the model cannot reach.

The paper measured a harness. Google's report is about the runner. According to the report, DUSTMAKER detects when it is running in CI/CD. "If confirmed, it extracts OIDC tokens from the process memory of GitHub Actions runners," the report says [13]. A harness check covers the tool calls an agent proposes [4], and a token read out of runner memory does not arrive as one. The commentary's author puts the ceiling on the job. "Permission has to be a property of the job, fixed before the agent reads a single untrusted byte," the author wrote [25]. The author is specific about the config lines: an agent job that holds id-token: write can publish, and one that holds actions: write can delete the logs that would show it did [17]. The report says DUSTMAKER creates pipeline tasks under AI-themed names such as 'Copilot Setup' and "issues automated API calls to delete the workflow execution logs" [16].

Rules written for the agent sit in files the attacker can reach. A rule in CLAUDE.md that says "do not push to main" is, the author wrote, "a request to a reader, and the attacker can write to that file too" [18]. DUSTMAKER writes there. It drops or modifies files in .claude/, .vscode/ and .cursor/ and uses them "to instruct the AI assistant to run arbitrary commands or scripts" [15]. The attackers keep their agent rules in markdown too: one actor in the report ran a credential-harvesting campaign in under six hours using "preconfigured markdown instruction sets as operational playbooks" [20].

Provenance checks inherit the same problem. The author wrote that "a signature is evidence about who held the token, and when the token came out of your own runner the attestation is valid and tells you nothing" [19].

Putting the ceiling on the CI job is the commentator's conclusion from reading the two documents together [1]. The author could not find a count of repositories DUSTMAKER hit or what the stolen tokens were scoped to when used, and the report's magnitude, "Thousands of third-party credentials," comes with an unnamed victim [21]. The commentator has not reproduced either document [22]. "I cannot tell you how likely this is to hit a given pipeline, only that it works," the author wrote [26].

What to watch

  • A count from GTIG or the affected registries of repositories hit by DUSTMAKER, and the scopes the stolen OIDC tokens carried when used.
  • A reproduction of the CapScope results beyond five Python repair tasks on the Pi agent.
  • The paper's account of the 3 CapScope runs where the injected effect still executed, and which capability each one used.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories