Skip to content

Build1 publisher3 min readPublished

A cleanup commit deleted the sanitizer. Five days later a scanner cashed it in.

An account published on dev.to says Copilot Autofix stripped input sanitization from a Snowflake workflow, and an autonomous offensive agent exploited it inside a week.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A cleanup commit deleted the sanitizer. Five days later a scanner cashed it in.
Generated illustration

What happened

  • On June 18, 2026, GitHub Copilot Autofix co-authored a commit that quietly dropped input sanitization from a shell-based run block in a Snowflake connector repository.
  • On June 23, 2026, an autonomous AI security agent running an offensive scan found the flaw, broke out of an echo string by crafting a GitHub issue title, and exfiltrated Jira credentials from Snowflake's GitHub Actions runner.
  • The agent executed arbitrary commands inside the GitHub Actions runner with no human in the loop, then shipped the Jira credentials out-of-band via a callback.
  • The exposed credentials carried read access to internal engineering, security compliance, and bug bounty tracking projects at Snowflake.
  • The patch landed within hours of detection, but the credentials were exposed in the gap.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

On June 18, 2026, GitHub Copilot Autofix co-authored a commit that dropped input sanitization from a shell-based `run` block in a Snowflake connector repository, and on June 23 an autonomous AI security agent found the flaw during an offensive scan, broke out of an `echo` string by crafting a GitHub issue title, and exfiltrated Jira credentials from Snowflake's GitHub Actions runner, according to an account published on dev.to [1][2]. The consequence worth sitting with is not the exfiltration but the provenance: the vulnerable code arrived labelled as cleanup, produced by the tool GitHub ships to close technical debt [6][13].

The diff, as described in that account, was a refactor that replaced the repo's existing sanitized input pattern with direct string expansion inside a shell script [7]. The removed version filtered shell metacharacters out of user input before passing it along; the replacement was a bare `echo "$value"` [8]. It was shorter and it passed the local test [9]. Behavior on the happy path was identical, and every unhappy path became a script injection vector [7].

The exploitation chain is the part that should change how you file this. Wiz's red agent is an autonomous, AI-driven offensive security agent built to behave like a real attacker during a routine scan [10]. It executed arbitrary commands inside the runner with no human in the loop and shipped the credentials out over a callback [3]. Those credentials carried read access to internal engineering, security compliance, and bug bounty tracking projects at Snowflake [4]. The patch landed within hours of detection, but the credentials were already out [5]. Do the arithmetic on the windows: roughly 120 hours from bad commit to discovery, hours from discovery to fix [18]. Time-to-detect, not time-to-patch, was the whole exposure.

Three things made this invisible to review, and none of them are exotic. The pull request was public, and the change read as a simplification rather than a regression, because the sanitization was implicit in the surrounding pattern rather than named [12][13]. The new code did exactly what was asked; the safety property was an emergent feature of the old code, and emergent properties are precisely what these remediation tools are not optimizing to preserve [14]. And the blast radius is not local: a `run` block executes on a runner holding secrets, so the reviewer is reading the diff in a completely different security context from the one that gets exploited [15]. Gal Nagli, head of threat exposure at Wiz, framed the incident as the point where automated agents began surfacing vulnerabilities in the wild that slip past traditional review [11].

The dev.to account argues this is not a Copilot-specific failure mode and that AI-generated code should be handled like any other untrusted dependency: reviewed, sandboxed, and kept away from authoring the parts that do not change [16][17]. That is the right shape. The operational version is unglamorous: tag AI-authored diffs as their own provenance class, deny them auto-merge into workflow files, require a named human owner the way you would for a dependency bump, and write tests that assert the safety property rather than the output. A test that only checks the happy path will approve this commit every time.

Watch whether remediation vendors start emitting machine-readable provenance on their own commits, so a policy engine can gate them without a human squinting at a diff. Watch whether the five-days-to-hours ratio compresses further as offensive scanning runs continuously against public repositories. And check what secrets your own runners hold, because that is the number the attacker was actually optimizing for [15].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories