Build1 distinct publisher3 min readPublished
The benchmark hands an agent an input that already crashes a program and asks it to escalate to file access or code execution. Claude Mythos Preview managed 157 of 898 instances, with model safeguards disabled.
The Engineer · Build desk

invest
Z.ai's 0.7-point CyberGym lead is a self-graded number on a model that is not yet open1 distinct publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
invest
Thomson Reuters trades Claude for a Qwen derivative it cannot let customers audit1 distinct publisher
leadership
Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist1 distinct publisher
Compiled by The EngineerSomething wrong?How this is made
Prior cybersecurity benchmarks graded vulnerability reproduction, patch generation and capture-the-flag puzzles, and frontier models already score well on many of them [11]. Those tasks mostly reward reading source. ExploitGym grades runtime behaviour instead: heap metadata, stack frames, virtual memory mappings, instruction-level control flow, register state, and inputs that satisfy tight constraints [10]. The task is escalation, from an input that crashes something to unauthorized file access or code execution [8], sustained over a long horizon while the process state moves under the agent [9]. A model that writes a correct patch has demonstrated nothing about whether it can hold a heap in a known shape across a hundred allocations.
Then the arithmetic. 157 working exploits out of 898 instances is 17.5 percent [13]. GPT-5.5's 120 is 13.4 percent [14], which puts Claude Mythos Preview ahead by about 1.3 times on raw count [15]. Those percentages only mean what they look like under conditions the abstract does not state. Both configurations would have to have been run against the full 898. An instance would have to be a distinct vulnerability rather than a vulnerability crossed with a protection setting, and since the authors vary the protections applied to each instance [5], one bug can plausibly appear several times, which would make the per-bug rate higher than the per-instance rate. And the reported leaders are described as configurations, not models [4], so the number grades a scaffold as much as a set of weights. Swap the harness and the count moves.
The bigger read-across problem is the test condition. Main experiments ran under trusted-access programs with safeguards disabled, explicitly to measure the capability boundary [7]. That answers the question the paper asks, but it describes what the weights can do rather than what a stranger with an API key gets out of the shipped product, which is not the number a defender needs for a threat model. Those are two different measurements and the abstract supplies one of them.
Same gap on mitigations. The paper reports that models retain non-trivial success rates with widely used defenses enabled, without a figure attached in the abstract [6]. Non-trivial is the least falsifiable number in the paper.
What survives all that caveating is still useful, and it lands on triage. Exploitation is what converts a severity argument into a demonstration, and the paper lists exactly that defensive use: assessing severity, prioritising patches, validating mitigations, alongside the acknowledgement that it lowers the expertise needed for offense [12]. If your process defers fixes on the strength of "crashes, probably not exploitable", this is the yardstick that eventually prices that habit.
The design choice worth copying is the one that gets least attention in the abstract. Varying the protections per instance and shipping every configuration as a reproducible container [3][5] turns a leaderboard question into a measurement of how much each mitigation costs an agent that can already exploit. That is a benchmark you can re-run against your own build flags instead of taking on faith.
Ranked by verification strength, evidence, and original report placement.
ExploitGym comprises 898 instances sourced from real-world vulnerabilities across three domains: userspace programs, Google's V8 JavaScript engine, and the Linux kernel.
The paper states that even with widely used defenses enabled, models retain non-trivial success rates; the abstract attaches no figure to that residual rate.
ExploitGym tasks agents with taking a program input that triggers a vulnerability and progressively extending it into a working exploit.
All ExploitGym configurations are packaged in reproducible containerized environments.
The strongest configurations evaluated were Anthropic's Claude Mythos Preview and OpenAI's GPT-5.5, which produced working exploits for 157 and 120 instances respectively.
The authors vary the security protections applied to each instance in order to isolate their impact on agent performance.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One preprint, its own scorekeeper
The numbers are precise and the methodology is unusually falsifiable — flag retrieval plus a judge that checks the exploit used the intended bug — but every figure traces to the team that built the benchmark, and nobody outside has run it. The available text also stops mid-sentence exactly where the scale-and-diversity detail would have been, so the 898 has no per-domain breakdown behind it and the defended-configuration results have no numbers at all.
Author-run only
What exists is an artifact and a self-run scoreboard: containers, 898 instances, two model scores. There is no third party using the benchmark, no vendor citing it in a system card, no CI pipeline or security team reporting results from it. Containerization makes uptake plausible; nothing in this reporting shows it happening.
Ceiling measured, risk generalised
The paper closes on 'growing cybersecurity risks', and the conditions behind that phrase are favourable to the agent in three ways at once: safeguards off, protections varied downward in the main runs, and the vulnerability already handed over as a crashing input. Strip that back and the best configuration finishes 157 of 898. The overstatement is modest rather than egregious because the authors disclose every one of those caveats themselves — and the unquantified 'non-trivial success rates with defenses enabled' is the one place where the rhetoric outruns the arithmetic entirely.
Gap-staking, vendor-dependent
Two incentives pull the same direction here. The paper stakes a claim to being the first comprehensive exploitation benchmark, which rewards framing the existing literature as a gap and the new capability as consequential. And the headline runs depended on vendor trusted-access programs with safeguards disabled — access that flows to researchers whose framing labs find useful. Nothing in this suggests bad faith; it does mean the risk narrative and the novelty narrative are both load-carrying for the authors.
Direction firm, detail missing
We are confident about what the paper says and how it grades, less confident about what the scores mean. A single self-reported source, a text that truncates before the dataset detail, no human baseline, no attempt budget, and no defended-configuration figures put a ceiling on this. The direction — frontier agents can escalate a supplied crash into code execution a meaningful minority of the time — looks solid; the magnitude is one lab's measurement.