Security1 publisher3 min readPublished
CrowdStrike's agent containment stack: seven layers, and escape treated as expected
The company says none of its offensive agents have broken out of their sandboxes so far. The useful part is the threat model, not the clean record.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- CrowdStrike says it conducts extensive red-team testing of agentic systems using diverse models, tools, and adversarial evaluation harnesses designed to probe for containment failures.
- CrowdStrike states that to date, none of its offensive agents have escaped their intended sandbox boundaries.
- CrowdStrike describes a secure-by-design, defense-in-depth approach built on seven independent control layers, each designed under the assumption that others may fail, with no single control trusted to contain the agents on its own.
- Layers 1 to 4 are containment boundaries, nested so an action must traverse outward through all of them to escape.
- Layers 5 and 6 are channel gates, described as peers, each policing a distinct action pathway.
Compiled by The WatchSomething wrong?How this is made
Why it matters
CrowdStrike has published the framework it uses to contain AI agents that autonomously run commands and call tools, describing seven independent control layers and stating that none of its offensive agents have escaped their intended sandbox boundaries to date [3][2]. The claim of a clean record is unverifiable from outside the company; the framing around it is the part worth copying, because it treats agent escape as an expected failure mode of a running system rather than a research curiosity.
The threat model is the load-bearing piece. CrowdStrike says it assumes the agent is capable and persistent, and treats it the way a security engineer treats untrusted code: "Contain by construction, not by trust" [7]. That assumption is applied regardless of where the unwanted behavior comes from, whether the model itself, adversarial content ingested through prompt injection, intent misalignment, or a defect in the harness [8]. It also assumes the agent may act through any available channel, including a shell, tool call, or network request, which is the stated reason each enforcement point exists independently [9]. This is ordinary untrusted-workload engineering, applied to a component that vendors more often describe as a colleague.
The architecture is grouped rather than itemised in the published post. Layers 1 through 4 are containment boundaries, nested so an action has to traverse outward through all of them to get out [4]. Layers 5 and 6 are channel gates, peers rather than a stack, each policing a distinct action pathway [5]. Layer 7 is cross-cutting and bounds volume at every boundary above it [6]. Four plus two plus one accounts for all seven, so every layer is characterised by role and none is named individually in the excerpt [17]. Note also that the post lists at least three action channels but only two channel gates, and does not say which gate covers which channel [18]. Teams trying to reproduce this will be reconstructing the specifics themselves.
Two details matter more than the layer count. First, CrowdStrike says the configuration used for each experiment depends on the agent, task, environment, access, and associated risks, and that for any given experiment some layers may be inactive [12]. A seven-layer architecture in which layers can be switched off is really a policy question about which configuration maps to which risk tier, and that mapping is where the mistakes will live. Second, containment here means preventing actions that are not intended or appropriate for the agent on a given task; in adversary emulation, breakouts and exfiltration may be deliberately in scope, and the controls are not meant to interfere with those [13]. The boundary is intent-relative, so the document defining scope becomes a control surface in its own right.
The context is not hypothetical. CrowdStrike says it reviewed public materials on documented incidents involving agent sandbox escapes, and designed its approach to contain those classes of behavior [15]. It also says the architecture applies to any harness where agents execute commands or invoke tools, across code analysis, software builds, security testing, data processing, and report generation [10], and that all the capabilities described are implemented and running [14]. Its own agents are used for vulnerability discovery, exploit development, post-exploitation, capture the flag, red-teaming, and adversary emulation [11].
Watch whether the layer-by-layer detail is published, and whether near-misses and misconfigurations are reported with the same enthusiasm as the zero-escape figure. A negative result with no stated denominator, no run counts, and no failure taxonomy is a starting position, not evidence.