Skip to content

Product1 publisher3 min readPublished

The allowlist read the command name, not what the command would do

CVE-2026-22708 let injected text rewrite a Cursor agent's environment, so an approved "git branch" ran something else. It worked with an empty allowlist too.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • On January 14, 2026, researchers at Pillar Security disclosed CVE-2026-22708, a flaw in Cursor: when the agent ran in Auto-Run Mode with an allowlist enabled, a handful of shell built-ins executed without appearing in that allowlist and without asking for approval.
  • Most teams running a coding agent today have some version of a list of commands the agent may run without asking, and the assumption underneath it is that anything dangerous will show up as a prompt you can refuse.
  • Programs read settings from their environment when they start up: Git checks one called PAGER to work out which program displays its output, and Python checks one called PYTHONWARNINGS.
  • The commands that change those environment settings, named in Pillar's research as export, typeset and declare, are shell built-ins rather than programs sitting on disk, and the checker was looking for programs on disk, so they went through without being surfaced.
  • The attack is two lines: export PAGER="open -a Calculator", which runs silently and is never surfaced for approval, followed by git branch, which the developer is asked about and approves.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

On January 14, 2026, researchers at Pillar Security disclosed CVE-2026-22708, a flaw in Cursor: with the agent in Auto-Run Mode and an allowlist enabled, a handful of shell built-ins executed without appearing in that allowlist and without asking for approval [1]. That matters because a list of commands the agent may run unprompted is the safety model most teams running a coding agent have today, resting on the assumption that anything dangerous surfaces as a prompt you can refuse [2].

The mechanics are short. Programs read settings from their environment at startup, and git consults one called PAGER to decide which program displays its output, while Python consults PYTHONWARNINGS [3]. The commands that change those settings, which Pillar's research names as export, typeset and declare, are shell built-ins rather than programs sitting on disk, and the checker was looking for programs on disk [4]. So the whole attack is two lines: set PAGER to the attacker's command, which runs silently, then let the developer approve git branch, which they will [5]. Git looks up PAGER to work out how to show the branch list, finds the payload there, and runs that instead [6].

No memory corruption was involved and no permission was escalated, according to Docker's account of the disclosure [7]. The developer saw an accurate prompt, approved a command that was genuinely harmless, and got arbitrary code execution, because the meaning of the command had been changed a minute earlier by something they were never shown [7]. Anything that can get text in front of the agent, a README, a dependency, an issue comment, is enough to change the variable [8].

The detail that settles the argument about tuning these lists: Pillar notes the attack still worked with a completely empty allowlist, the most restrictive setting on offer [9]. An allowlist with zero entries and an allowlist with ninety entries produce the same outcome here, so no configuration of the mechanism is a mitigation [10]. Cursor rated the flaw High and patched it in version 2.3 [11], and its documentation now describes the allowlist as best-effort and warns that bypasses are possible [12]. Pillar's position is more direct: hand agents full command execution inside an isolated environment, and deprecate allowlists altogether [13].

None of the underlying trick is new. Pillar's write-up points back to Elttam's 2020 research showing environment variables could be turned into code execution [14], which sat unremarked for six years [15] until agents started approving commands on a developer's behalf. The structural problem, as Docker's series frames it, is that the agent runs as you, with your filesystem permissions and your credentials, and nothing sits between the model's decision and the shell's execution [16]. Docker sells the conclusion it wants here, containment at the boundary rather than at the command line [17], and its own post concedes there are two things its sandboxes do not contain in this case [18], which the excerpt we have does not enumerate.

Watch whether other agent vendors follow Cursor in reclassifying allowlists as best-effort rather than a control, since that is a documentation change with procurement consequences [12]. Watch also whether approval moves off the laptop: Docker's pitch is that kits, organisation policy and audit logs cover what a per-laptop allowlist misses [19], and any team relying on per-developer configuration has no record of what was approved or why.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories