Skip to content

Build1 publisher3 min readPublished

EDR sees the file write, not the reason: the case for an agent-native detection layer

A framework running at Uber for over ten months argues endpoint tooling is structurally blind to agent reasoning. On the authors' own benchmark it still misses a third of attacks.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying EDR sees the file write, not the reason: the case for an agent-native detection layer
Generated illustration

What happened

  • The Agentic AI Detection and Response (ADR) system is presented by authors including Chenning Li, Pan Hu, Justin Xu, Mohammad Alizadeh, Pengyu Zhang and Ming Zhang, with a copyright line for the Proceedings of the MLSys Conference (Industry Track), Bellevue, WA, USA, 2026; the authors describe it as the first large-scale, production-proven enterprise framework for securing AI agents operating through the Model Context Protocol (MCP).
  • The paper identifies limited observability as a persistent challenge: existing Endpoint Detection and Response (EDR) tools see file writes and network calls but not the agent reasoning, prompts, or causal chains linking intent to execution.
  • MCP is described as a standardized interface for connecting large language models to external tools and data sources, used by enterprises to deploy agents that analyze documents, modify infrastructure, generate code and interact with internal systems.
  • The paper states that agents can be manipulated through natural language, exploited via compromised MCP servers, or coerced into executing unsafe commands and exfiltrating sensitive data.
  • The paper cites insufficient robustness as a second challenge: static defenses constrained by pre-defined rules and policies fail to generalize across diverse attack techniques and enterprise contexts, from prompt injection to tool manipulation to credential exfiltration.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A team of Uber-affiliated authors and academic collaborators has published the design and production record of ADR, a detection and response system for AI agents operating through the Model Context Protocol, carrying a copyright line for the MLSys Conference Industry Track in Bellevue, Washington, 2026 [1]. Their central claim is structural rather than incremental: endpoint detection and response tools can observe a file write or a network call, but not the user prompt, the agent's reasoning steps, or the causal chain that links intent to that execution [4].

That is a real gap, and it is not one you close with tuning. MCP is a standardised interface for connecting large language models to external tools and data sources [5], and the paper argues it opens an attack surface where agents can be manipulated through natural language, exploited via compromised MCP servers, or coerced into running unsafe commands and exfiltrating data [6]. An EDR alert on a suspicious write tells you a process misbehaved; it does not tell you whether a poisoned tool description talked the agent into it. The authors also fault static guardrails and rule-based policy checkers for failing to generalise across attack techniques and enterprise contexts [7], and flag severe class imbalance as a further obstacle [2].

The third constraint is the one operators will feel first: the paper states that LLM-based inference is prohibitively expensive at scale [8]. ADR's answer is a three-part split. A lightweight endpoint sensor reconstructs high-fidelity agentic telemetry [9]; an "Explorer" component does pre-deployment red teaming and hard-example generation [10]; and the detector runs two tiers, fast triage in front of context-aware reasoning [11]. In other words, you do not get to put a reasoning model on every session, so you build a cheap filter and reserve the expensive judgement for what survives it.

The deployment numbers are the most interesting part, because they are operational rather than benchmark. ADR has run for more than ten months across over 7,200 unique hosts, processing more than 10,000 agent sessions daily [12]. It surfaced hundreds of credential exposures across 26 categories, and the authors report a shift-left prevention layer at 97.2 percent precision with 206 detected credentials [13]. Credential leakage, not exotic agent hijacking, is what the telemetry actually found at volume.

The benchmark results are more mixed than the framing suggests. On ADR-Bench, which the same team introduces with 302 tasks, 17 techniques and 133 MCP servers, ADR reports zero false positives while detecting 67 percent of attacks, which the paper says beats ALRPHFS, GuardAgent and LlamaFirewall "by 2-4 in F1-score" [3]. Zero false positives is worth something; missing roughly a third of attacks on your own benchmark is worth remembering [16]. On the public AgentDojo prompt injection benchmark it detects all attacks with three false alarms across 93 tasks [14], about 3.2 percent of tasks [17].

Two caveats. These are self-reported figures from a single enterprise deployment, and the ADR-Bench scoreboard was built by the party being scored [3]. The version available is an arXiv rendering carrying automated warnings that its page layout violates the ICML style [15].

Worth watching: whether anyone outside Uber reproduces the 67 percent figure, whether the 97.2 percent precision on credential prevention holds as agent traffic grows past 10,000 daily sessions, and whether ADR-Bench is actually released in a form competitors can run.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories