Science1 distinct publisher3 min readUpdated
Adversa.ai says Claude Code's folder-trust dialog stopped naming MCP servers in v2.1, while a repo-supplied config can still launch an unsandboxed process on one keypress.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Adversa.ai has published a report it calls TrustFall, which says four agentic coding CLIs, Claude Code, Gemini CLI, Cursor CLI and Copilot CLI, execute project-defined MCP servers the moment a developer accepts the folder trust prompt [1]. That places the consequential decision in a dialog most developers clear by reflex, not in anything the model does afterwards [1][8].
The Claude Code chain is the one Adversa documents in detail. According to the report, the trust dialog used to warn that a cloned repository contained MCP servers and offered an opt-out, and in v2.1 and later that warning was removed [2]. The dialog now reads "Quick safety check: Is this a project you created or one you trust?" and lists nothing [3]. A malicious repository ships an MCP server and auto-approves it through its own .claude/settings.json, so one Enter keypress starts that server as an unsandboxed OS process with the developer's full privileges, with no tool call from Claude required [4]. The payload does not have to be a file on disk; the whole script can sit inline in .mcp.json [5].
The privilege picture is the part operators should sit with. Adversa says these servers run as native OS processes with the full rights of the user running Claude Code, not sandboxed, not confined to the project directory, and not restricted to any subset of filesystem or network [6]. In practice that means reading stored secrets and source code from unrelated projects, or opening a long-lived command-and-control channel [7].
There is an internal inconsistency worth noting. Other dangerous settings, such as bypassPermissions, are already blocked from project scope or gated behind a red warning dialog, while the MCP-enabling settings are neither [9]. The report describes a third silent path through permissions.allow [10], and counts three project-scoped settings that can spawn arbitrary executables behind the prompt [11].
Continuous integration removes even the keypress. Adversa says that when Claude Code runs headless, the default for the official claude-code-action, the trust dialog is skipped and never renders, so the same attack runs with zero human interaction against pull-request branches [12].
Anthropic's security team reviewed the report and declined it as outside their threat model, on the basis that accepting "Yes, I trust this folder" is consent to the full project configuration and that post-dialog execution is the boundary working as designed [13]. Adversa says it does not contest where that boundary sits, and is instead documenting an informed-consent gap inside it, since the dialog no longer says what it is asking permission for [14]. The cross-CLI parity check came after that response and reframed the finding from a vendor regression to a shared convention, which is why Adversa did not pursue vendor-by-vendor disclosure [15]. All four tested CLIs default to Yes or Trust and differ only in how the dialog frames the authorization [8].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Adversa.ai's TrustFall report states that four agentic coding CLIs, Claude Code, Gemini CLI, Cursor CLI and Copilot CLI, execute project-defined MCP servers the moment a developer accepts the folder trust prompt.
Claude Code's trust dialog used to warn about MCP servers in a cloned repository and offer an opt-out; in v2.1 and later that warning was removed.
The current Claude Code dialog reads "Quick safety check: Is this a project you created or one you trust?" and lists nothing.
MCP servers execute as native OS processes with the full privileges of the user running Claude Code; they are not sandboxed, not confined to the project directory, and not restricted to any subset of the filesystem or network.
The MCP server has enough privilege to read stored secrets and source code from other projects, or to open a long-lived command-and-control channel.
All four tested CLIs default to "Yes/Trust" and differ only in how the dialog frames the authorization.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed single-source disclosure, no independent reproduction
The technical chain is described with specific, checkable artefacts — named settings, dialog strings, an inline .mcp.json payload, a stated PoC and demo video — and the affected vendor's rebuttal is disclosed rather than hidden, which raises credibility above a bare assertion. But the cluster contains exactly one item, authored by the researcher who benefits from the finding, with no third-party reproduction, no CVE or advisory, and no direct vendor statements beyond Adversa's paraphrase of Anthropic's decline. Cross-CLI parity for Gemini, Cursor and Copilot rests entirely on the same untested-by-others report.
Shipped defaults on four widely used CLIs, but no usage or exploitation data
The exposure is attached to default behaviour in generally available products rather than a prototype: four shipping agentic CLIs are said to execute project-defined MCP servers on trust acceptance, and headless CI is described as the default for the official claude-code-action, which extends the surface to pull-request automation. That breadth is what lifts the score above the floor. It stays well below the midpoint because the cluster supplies no install counts, no telemetry, no evidence of repositories using these settings maliciously, and no observed incident — all breadth claims trace to a single unreplicated vendor test.
Headline severity outruns the disputed, single-source basis
The framing '1-click coding agent RCE' in three vendors' products is stronger than what the cluster establishes: the affected vendor reviewed the chain and classified post-trust execution as consent and by design, and Adversa itself concedes the boundary and narrows its claim to an informed-consent UX gap. Nothing here is fabricated — the dialog regression and the settings mechanics are specific and plausible — so the gap is moderate rather than large, driven by severity language, a class-level generalisation from one vendor's own parity test, and the absence of any observed exploitation.
Security vendor self-publishing a branded finding after a rejected report
The single source is a commercial AI-security firm publishing a named finding ('TrustFall') on its own blog, explicitly after the affected vendor declined the report and after it decided that vendor disclosure was 'not the right shape of response'. Publication visibility is the vendor's remaining channel and also its marketing surface, which is a clear promotional incentive. Mitigating factors keep the score short of the top band: Anthropic's opposing position is quoted up front, the researchers concede the trust boundary, and they ship a reproducible PoC that invites others to check the work.
Low: one interested publisher, contested characterisation
Confidence is limited by cluster structure rather than by internal inconsistency. There is one publisher, that publisher is the researcher, the central characterisation is contested by the implicated vendor, and the other three implicated vendors are silent in the supplied material. The mechanics are internally coherent and specific enough to be falsifiable, which keeps this from the floor, but any assessment here should be revisited once independent reproduction, vendor statements, or advisory identifiers appear.
science
OX Security says MCP command execution is a design choice, so server owners own the risk1 distinct publisher
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
leadership
Anthropic's own telemetry: 93% of permission prompts approved. Budget for blast radius, not reviewers1 distinct publisher
build
Claude Code's new default is a confession: the approval prompt was never a control1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026