Build2 publishers2 min readPublished
GPT-5.6 Sol Ultra turned public V8 patch commits into a working Chrome exploit for $1,597
Hacktron reports that GPT-5.6 Sol Ultra built a complete Chrome/V8 exploit chain in a controlled test for $1,596.89 of model compute. The benchmark handed the model the source tree and the public security-fix commits, so the figure covers only the compute for one lab task.
The Engineer · Build desk

What happened
- GPT-5.6 Sol Medium and Grok 4.5 reached the information-leak stage but never built arbitrary read/write inside the sandbox, got stuck in dead ends, and Hacktron stopped both runs.
- In the lab the chain reached code execution, demonstrated by launching the Calculator app on the target machine.
- Hacktron first published the benchmark on July 14, 2026, and recirculated the result in a September 29 post on X without reporting any new vulnerability.
- Hacktron says the run is not evidence that the model compromised any deployed Chrome install, or that the chain works against a current, patched browser.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint At $1,596.89 the figure buys compute for one run and nothing more, since the benchmark supplied the source tree and the fix commits that normally take the most research.
- precedent The model worked from public security-fix commits, so the result speaks to n-day exploitation, the window between a patch shipping and teams installing it.
- contradiction Hacktron frames the run as the end of exploit development at inference scale, while runtimewire reads the same result as a sharp advance in one controlled task.
A fix commit tells you what was wrong and where. Hacktron gave all three models exactly that, together with the V8 source tree at version 14.9.207.35, the engine in Chrome 149.0.7827.201, and a sandbox-enabled d8 shell [3][13]. The job was to walk from a published patch back to a working exploit for the build before it [14].
Sol Ultra chained five stages [18]. It opened with a Maglev type confusion, where the compiler inlined ArrayIterator.prototype.next() but skipped a map re-check, forged a JSArray header to build a 4GB read/write primitive, expanded that to the full 1TB sandbox through DataView metadata, leaked native addresses through a signed-offset bug in String::VisitFlat, and finished with a use-after-free during background Wasm compilation [18]. The last stage handed over program-counter control [14].
Over three days, the agent managed its own context. The root agent spawned 74 subagents, and Pedhapati reported that roughly 70% of the investigation ran inside them [15]. The root compacted its context 33 times, the full agent tree 70 [16]. On average a root compaction cut active input from 262,869 tokens to 18,911, a 92.67% reduction [16].
The benchmark supplied the hardest inputs itself, the source tree and the commits that pointed to each bug [7]. The run processed 2.096 billion tokens across 14,062 requests [6]. At $1,596.89 that is about $0.76 per million tokens, blended across input and output [1].
Pedhapati, a former senior security researcher at Cure53 [8], titled his write-up around the claim that "exploit development as we know it is dead for anyone who can throw inference-scale compute at a model at GPT-5.6's level or above" [19]. Hacktron sells the workflow this test resembles. Its product does source-aware code review that validates vulnerabilities in pull requests, plus white-box penetration tests [9]. The company disclosed a $2.9 million pre-seed round in May and about $240,000 in revenue over nine months, both company-reported [10]. runtimewire, flagging the distance between the controlled experiment and the X post that recast the cost as a warning, called the run a sharp advance in one controlled task [11][12].
What to watch
- Whether anyone reproduces the chain without being handed the source tree and the fix commits.
- Whether Hacktron's final_poc.js still reaches code execution against current patched Chrome, given the test targeted a pinned pre-patch V8 build.
- Independent verification of the compute cost, token counts and run time, all currently company-reported.