Security1 publisher2 min readPublished
Claude Opus 4.6 cost 40 times as much as Gemini 3 Flash per run for 11 more points of pen-test coverage
Ridge Security's 96-test pen-test benchmark had Claude Opus 4.6 reach 63% coverage at $217 a run, against 52% for Gemini 3 Flash at about $5.42. Ridge argues tooling matters more than the model, yet its published runs held the tooling fixed and varied the model.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Grok 4.5 posted the highest coverage of the models in Ridge's benchmark, at 77%.
- GPT-OSS-120B was the most efficient model tested, producing 16.9 findings per million tokens at $2.32 a run.
- Ridge concluded a harness is a must-have layer in which the model reasons, the harness controls execution, and independent validation confirms each finding.
- Ridge describes the work as the first public benchmark comparing multiple leading LLMs in autonomous penetration-testing workflows.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- contradiction Reported coverage spread at least 25 points between models inside one harness, so Ridge's own data show model choice moving results by a quarter of the scale.
- cost Each extra point of coverage Claude Opus 4.6 adds over Gemini 3 Flash costs about $19 per run, and that premium recurs with every test a team schedules.
- decision Hadrian's Rogier Fischer told buyers to judge agents on false-positive rates, repeatability and cost per validated finding; Ridge published coverage and cost per run, so those comparisons fall to the buyer.
This is a lab result. Ridge pointed the agents at intentionally vulnerable environments and scored whether each model could carry a test through reconnaissance, hypothesis testing, payload adaptation, exploitation and verification without stalling [1][3]. Ninety-six tests across eight models works out to 12 per model if they were split evenly [5].
The ReversingLabs writeup does not report any run without Ridge's harness or through a different one [1]. Its lead finding is still that performance in offensive security depends far more on the system around the model than on the model itself [18]. Lydia Zhang, Ridge's president, said: "Our research shows that autonomous offensive security is a systems problem. The model needs an architecture around it that can manage execution, adapt to what it discovers, and verify that a finding is real." [8]
Rogier Fischer, chief executive of Hadrian, described a penetration test as a long-horizon task that depends on state tracking, reliable tool use and proof of every finding [20]. "A leaderboard ranking is not a procurement criterion," he said [11]. In his account, the harness supplies scope enforcement, memory, deterministic tooling, validation and an audit trail [13]. "A frontier model without a harness is a brilliant intern with root access and no supervision," Fischer said [14].
Li Zhao, a principal strategic services consultant at Black Duck Software, said: "A well-engineered agent running on a smaller or more cost-effective model can often outperform a frontier model when the surrounding architecture, tooling, and workflows are optimized for offensive security operations." [15] Seemant Sehgal, founder and chief executive of BreachLock, said Ridge's findings confirm what practitioners have known for some time [19].
Ridge's researchers found that frontier models deliver stronger coverage at substantially higher cost per run [17]. Per percentage point of coverage, a Claude Opus 4.6 run cost about $3.44 and a Gemini 3 Flash run about 10 cents, roughly 33 times apart [4][7]. An Opus run cost about 94 times as much as a GPT-OSS-120B run [6].
What to watch
- Ridge publishing the full table for all eight models, including Grok 4.5's cost per run and how coverage was scored.
- A rerun of the same targets through a second harness, or with no harness, to measure the effect Ridge attributes to tooling.
- Per-model false-positive and repeatability figures from Ridge or another tester, the measures Fischer says buyers should use.