Security1 distinct publisher3 min readPublished
The three escapes used a kernel bug disclosed weeks earlier, a libslirp flaw Debian 12 has yet to fix, and 0-days the agent found itself after QEMU was rebuilt from upstream. It is one researcher's account.
The Watch · Security desk

security
A frontier model left a patched QEMU guest three ways. Two needed no new bugs.1 distinct publisher
build
qm start said nothing, the UI said stopped, and only host dmesg named the passthrough failure1 distinct publisher
leadership
Attackers reached a domain controller through one exposed SolarWinds helpdesk1 distinct publisher
build
A prebuilt orchestrator binary demanded GLIBC_2.32 from a node that ships 2.311 distinct publisher
Compiled by The WatchSomething wrong?How this is made
The libslirp chain stands out. libslirp is the library that gives a guest networking [11], so it is loaded precisely because you wanted the agent to reach the internet [17]. Debian 12 still ships version 4.7.0, vulnerable to CVE-2026-9539 [9]. On its own that is a patch-lag story. The agent then found a second libslirp bug-fix commit that carried no CVE at all, and combined the two into an arbitrary memory read and write in the host, after which escape is trivial [10]. Any inventory control keyed to CVE identifiers sees half of that chain.
There were three escapes in total [3]. The first used Januscape, CVE-2026-53359, in the host kernel, and it closed when the kernel went to the newest build Debian 12 offers [7][8]. The second worked on that updated host [9][10]. The third came after QEMU and its dependencies were rebuilt from the latest upstream source, and it was several 0-days [12]. Two of three paths survived a fully patched stack [19].
The operating mode matters more than the bug list. Hours running autonomously, backtracking off approaches that failed, pulling code and research papers, writing oracles, cutting its own minimal examples, aiming at a reusable and reliable exploit rather than a one-shot crash [13]. The researcher's contribution was walking to the machine and power-cycling it after each host hardlock [14]. That is the workload profile of a staffed vulnerability research effort.
What is public is a single first-person account from the person who ran the test [22]. There is no public exploit for Januscape, and the agent's version did not land cleanly on the first attempt [7]. The upstream 0-days come with no identifiers and no count beyond "several" in the material available [21]. The agent was also started outside the VM with SSH access into it, by design, because escape attempts crash the guest [4][5]; the author raises the cheating risk himself and says a final exploit could be audited, and the outcome of any such audit is not reported [20]. So far, no one outside Trail of Bits has reproduced the result.
The author's own assessment shifted over the course of the test: he started out doubting the model could pull this off at all, came to think it might manage it eventually, and ended up expecting it to happen fast [16][15]. His conclusion is that a mere VM should no longer be assumed to contain a sufficiently advanced agent. The design consequence is narrower than the slogan. A sandbox that earns its keep by sharing screen, files and network with the host [17][18] is worth exactly what QEMU's device and networking code is worth on the day you run it.
Ranked by verification strength, evidence, and original report placement.
Trail of Bits received preview access to GPT 5.6-Cyber as part of Patch the Planet, with the task of evaluating its cyber capabilities.
The target was a QEMU/KVM VM on the author's Linux dev machine, running Debian Linux 12 on AMD Zen3.
The task was framed as a CTF: the agent had to break out of the VM and read a flag file. It was started outside the VM, given SSH access into the VM, permitted to do anything inside it, and told to keep a research log.
The author started the agent on the host because escape attempts may lock or crash the VM during operation, judged the risk acceptable because nothing of importance was on the dev machine, and noted any final exploit could be audited for cheating.
About an hour after the agent was let loose, the host machine was hardlocked.
First escape: the agent found the host kernel vulnerable to Januscape (CVE-2026-53359), disclosed a few weeks earlier. There is no public exploit; the agent built one and tested it, but it did not land cleanly.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named CVEs, unnamed 0-days, one witness
Two of the three escape paths are specific enough to argue with — Januscape by CVE number, libslirp 4.7.0 by shipped version plus the fix commit it was chained against. The third, the one that mattered most because it survived an upstream rebuild, is "several 0-days" and the write-up stops one word into that section. No exploit code, no log excerpts, no second observer, and the researcher's own admission that the first attempt did not land cleanly all sit in the same document.
One dev box, one preview seat
Real-world footprint is a single Linux development machine and a single preview allocation handed to one firm. What has broader reach is the incidental finding: the libslirp and CUPS gaps belong to Debian oldstable as shipped, not to this researcher's configuration, so anyone on that base carries them. Nobody else in this reporting has run the exercise or changed how they isolate agents because of it.
Three in the headline, two and a half in the text
"Escaped three different times" is doing more work than the escapes support. The first exploit did not land cleanly and is scored a point on the author's confidence that revision would have fixed it; the third is asserted without identifiers in a section that never arrives. Only the libslirp chain is shown end to end. Against that, the conclusion — stop assuming a VM holds an agent — is stated with unusual restraint about its own basis, and the author volunteers the cheating risk he created. Overstated at the headline, honest in the body.
Preview seat in, capability story out
The pipeline runs both ways. A model vendor granted a security firm early access under Patch the Planet; the firm published a result that flatters the model's offensive reach and simultaneously argues that containing such agents is hard work of exactly the kind it sells. Neither party is served by a dull finding. The countervailing detail is that the post names its own weak spots — the failed first exploit, the host-side launch — which a pure capability advertisement would have trimmed.
Checkable in principle, unchecked in fact
Enough is nailed down that an operator can go look — the CVE numbers and the libslirp version are falsifiable this afternoon on any Debian 12 box. Everything about the agent's behaviour is not: no artifacts, no replication, no result from the cheating audit the author himself said was possible, and the decisive third escape trails off mid-section. Confidence sits at the midpoint because the verifiable half is verifiable and the load the thesis carries rests on the other half.