Skip to content

Security1 publisher2 min readPublished

Claude Opus 4.6 cost 40 times as much as Gemini 3 Flash per run for 11 more points of pen-test coverage

Ridge Security's 96-test pen-test benchmark had Claude Opus 4.6 reach 63% coverage at $217 a run, against 52% for Gemini 3 Flash at about $5.42. Ridge argues tooling matters more than the model, yet its published runs held the tooling fixed and varied the model.

The Watch · Security desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Claude Opus 4.6 cost 40 times as much as Gemini 3 Flash per run for 11 more points of pen-test coverage
Generated illustration

What happened

  • Grok 4.5 posted the highest coverage of the models in Ridge's benchmark, at 77%.
  • GPT-OSS-120B was the most efficient model tested, producing 16.9 findings per million tokens at $2.32 a run.
  • Ridge concluded a harness is a must-have layer in which the model reasons, the harness controls execution, and independent validation confirms each finding.
  • Ridge describes the work as the first public benchmark comparing multiple leading LLMs in autonomous penetration-testing workflows.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • contradiction Reported coverage spread at least 25 points between models inside one harness, so Ridge's own data show model choice moving results by a quarter of the scale.
  • cost Each extra point of coverage Claude Opus 4.6 adds over Gemini 3 Flash costs about $19 per run, and that premium recurs with every test a team schedules.
  • decision Hadrian's Rogier Fischer told buyers to judge agents on false-positive rates, repeatability and cost per validated finding; Ridge published coverage and cost per run, so those comparisons fall to the buyer.

This is a lab result. Ridge pointed the agents at intentionally vulnerable environments and scored whether each model could carry a test through reconnaissance, hypothesis testing, payload adaptation, exploitation and verification without stalling [1][3]. Ninety-six tests across eight models works out to 12 per model if they were split evenly [5].

The ReversingLabs writeup does not report any run without Ridge's harness or through a different one [1]. Its lead finding is still that performance in offensive security depends far more on the system around the model than on the model itself [18]. Lydia Zhang, Ridge's president, said: "Our research shows that autonomous offensive security is a systems problem. The model needs an architecture around it that can manage execution, adapt to what it discovers, and verify that a finding is real." [8]

Rogier Fischer, chief executive of Hadrian, described a penetration test as a long-horizon task that depends on state tracking, reliable tool use and proof of every finding [20]. "A leaderboard ranking is not a procurement criterion," he said [11]. In his account, the harness supplies scope enforcement, memory, deterministic tooling, validation and an audit trail [13]. "A frontier model without a harness is a brilliant intern with root access and no supervision," Fischer said [14].

Li Zhao, a principal strategic services consultant at Black Duck Software, said: "A well-engineered agent running on a smaller or more cost-effective model can often outperform a frontier model when the surrounding architecture, tooling, and workflows are optimized for offensive security operations." [15] Seemant Sehgal, founder and chief executive of BreachLock, said Ridge's findings confirm what practitioners have known for some time [19].

Ridge's researchers found that frontier models deliver stronger coverage at substantially higher cost per run [17]. Per percentage point of coverage, a Claude Opus 4.6 run cost about $3.44 and a Gemini 3 Flash run about 10 cents, roughly 33 times apart [4][7]. An Opus run cost about 94 times as much as a GPT-OSS-120B run [6].

What to watch

  • Ridge publishing the full table for all eight models, including Grok 4.5's cost per run and how coverage was scored.
  • A rerun of the same targets through a second harness, or with no harness, to measure the effect Ridge attributes to tooling.
  • Per-model false-positive and repeatability figures from Ridge or another tester, the measures Fischer says buyers should use.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories