Security1 distinct publisher2 min readPublished
Trail of Bits gave a preview model one job: escape the sandbox on its author's Debian 12 host. The route worth studying needed only a published CVE and an unlabelled upstream fix commit.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
The order of the three routes is the finding. Configuration one: a kernel bug disclosed a few weeks before the test, with no public exploit, so the agent wrote its own [5]. Configuration two, after a full update: a CVE in libslirp 4.7.0 that Debian 12 was still shipping unfixed [7], paired with an upstream bug-fix commit carrying no CVE, and together they produced arbitrary read and write on the host [8]. Configuration three, with QEMU and its dependency tree rebuilt from latest upstream: several 0-days [c4b]. Each hardening step changed which class of bug the agent needed, and changed nothing about whether it got out [20].
Two of the three routes therefore required no original vulnerability research [18]. That is the part that transfers to people who will never get preview access to this model. The libslirp path is patch latency converted into an exploit primitive: the fix existed upstream, the advisory existed, the shipped package did not follow. The author runs oldstable precisely because it changes slowly and still receives security updates, and notes that some of those updates do not actually arrive in the distribution [17]. He also writes that he did not know libslirp was in his stack [21]. It is there because guests need networking [9].
Diffing upstream commit histories for quietly fixed memory bugs, then working out what each one fixes, is tedious and open-ended work for a person. In this account the agent does it unattended for hours, abandoning approaches that fail, pulling code and research papers, writing oracles and building its own minimal examples in pursuit of a reusable exploit [10], while the human contribution reduces to power-cycling a hardlocked machine [11]. The capability on display is duration at low cost rather than insight.
Note also which resources the escapes used: the shared state, the file and screen paths, the network access that make a sandbox usable in the first place [16]. Remove them and the guest cannot do the build-and-debug work the agent was put there to do. That is why the author's conclusion lands on threat model rather than on virtualisation: treat a sufficiently capable agent as an advanced persistent threat, and stop assuming a mere VM contains it [12]. His own stated confidence moved from "no, absolutely not" through "maybe, after a while" to "yes, and quickly" over the course of the exercise [15]. This is one engineer, testing one host, with a model that is not generally available [1], so the transferable claim is the narrower one: the window between an advisory and a shipped package is now something an unsupervised process can work through in an afternoon.
Ranked by verification strength, evidence, and original report placement.
The challenge given to the model was to escape the QEMU/KVM VM the author normally uses for sandboxing; the target was a QEMU/KVM VM on his Linux dev machine, Debian Linux 12 on AMD Zen3.
The agent escaped the VM three different times.
After the author fully updated the host, the agent used disclosed bugs that had not yet reached package maintainers or were not classified as security bugs.
The agent operated autonomously for hours, backtracked from approaches that did not work, pulled code and research papers, wrote oracles, made its own minimal examples, and aimed for a reusable, reliable exploit with minimal handholding and prompting.
The author's main job during the exercise was to physically reboot the machine when the agent hardlocked the host kernel.
The author states plainly that you can no longer assume a mere VM will contain a sufficiently advanced AI agent, and that such agents should be treated as an advanced persistent threat.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-party account, no artifacts or replication
The libslirp route is specified concretely enough to check against distribution package versions, and the failure paths (CUPS blocked by AppArmor, unreliable hardware read oracle) are reported rather than hidden, which raises credibility. Against that: one publisher, one host, one operator, no exploit code or bug identifiers, the headline 0-day escape section truncated, and the first counted escape admittedly did not land. Nothing here has been independently reproduced.
One preview evaluation plus one documented distro exposure
Observable uptake is limited to a single vendor evaluation run disclosed on the vendor's blog, alongside a concrete distribution patch-lag exposure on the author's own host. There is no evidence of other labs, operators, or teams reproducing the escapes, changing sandbox architecture in response, or any usage, deployment or pricing disclosure.
Universal conclusion drawn from one unhardened host
The stated conclusion - that no mere VM will contain a capable agent, and such agents are an advanced persistent threat - generalises well beyond what one run on a Debian 12 oldstable box with mitigations=off, an unpatched CUPS package and a stale libslirp can establish; no hardened or network-isolated configuration was tested, and the 0-day escape that anchors the strongest reading is undocumented. The gap is positive but moderate rather than severe, because the libslirp chain is a genuine, verifiable result that arguably deserves more attention than it gets in the framing.
Security vendor publishing a vendor-granted capability demo
The author is a security consultancy that received preview model access under a named programme and then published a dramatic capability result; both parties benefit - the evaluator from demonstrating agent-security expertise and demand for containment work, the model provider from a credible third party attesting to frontier cyber capability. The post discloses the access relationship but not its commercial terms or any review arrangement, and no independent party checked the findings.
Credible operator detail, thin verification base
Moderate confidence: the technically checkable core (distro shipping a vulnerable libslirp, chaining a public CVE with an unmarked fix commit, shared-surface argument) is coherent and specific, and the account volunteers its own failures. Confidence is capped by the single-source cluster, absent artifacts, truncated 0-day section, and the incentive alignment between evaluator and model provider.
build
qm start said nothing, the UI said stopped, and only host dmesg named the passthrough failure1 distinct publisher
build
A 28-host Debian 12 cutover, and the 02:13 failure Ansible could not have prevented1 distinct publisher
security
Provenance's marker module let anyone with zero tokens claim admin over 82 live financial assets1 distinct publisher
security
Three Zoom annotation bugs made every screen share a two-way takeover path2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026