Build1 distinct publisher3 min readPublished
Aikido spent 11.7 billion tokens rediscovering 32 fresh CVEs with ten models, three attempts each. The number that should move a scanning budget is the marginal cost of the second and third pass.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Pooled recall is the size of a union. Three passes over the same 32 cases produce three sets of findings, and the score counts every case that turns up in at least one set [10][3]. That is why repetition reorders the table. DeepSeek V4 Pro 0813 returned 17 on its first pass and 28 pooled, while Grok 4.6 went from 24 to 26 because its runs kept covering the same ground, according to Aikido [6][7][14].
Price that union and the shape of the spend shows up. Three Pro passes cost about $295, so one pass is roughly $98 [4][1]. The first pass bought 17 findings at about $5.80 each [2]. The next two bought 11 more at about $17.90 each [3]. Repetition is worth buying, and it is worth buying at a rising unit price, which is what sampling a partly explored search space looks like when you write the arithmetic down.
Flash is the line worth arguing about. Three DeepSeek V4 Flash 0731 passes cost $108 and reached 24 [5]. That is 0.22 findings per dollar against Pro's 0.095, a 2.3x advantage on recall per dollar [4]. The same $295 would fund eight Flash passes instead of three Pro passes [5], and nothing in the published data says where the Flash curve flattens.
Aikido also says the Flash triple matched Grok's best individual pass for under a quarter of the cost [5]. Read strictly, one Grok pass then sits above $432 [6]. Read as a quarter of three Grok passes, it sits above $144 [6]. Per-model pass prices are not published, so the defensible form of that claim is the ratio.
The dataset caps how much of the ranking is signal. Seventeen of the 32 cases were found by every model at least once, and one case defeated all ten [13]. Fifteen cases therefore separate the field, and only 14 of those were winnable by anyone [7]. DeepSeek Pro's 28 against Grok's pooled 26 is a two-case gap inside that band [8].
What the dollar figures leave out is triage. Aikido reports that the open models caught up on recall at much cheaper rates while also producing the most false leads for the pipeline to reject [9]. A union pools findings and it pools noise, so whoever reads the queue pays for pass two and pass three a second time, in hours. The benchmark measures recall against an answer key and defines precision as the share of reported findings accepted [15]; production has no key, which means no stopping rule. Pass four costs the same $98 whether it adds a real bug or 30 turns of confident restatement.
Three settings decide whether any of this transfers. The agent got 30 turns and no internet access, inside a bounded version of the harness behind Aikido's own AI Code Analysis product [11]. The vendor sells the harness, which is a reason to read the setup section before the TL;DR. The 32 vulns were pulled from recently disclosed CVEs specifically to reduce the chance the models had seen the write-up or patch in training, with cases, prompts, tools and evaluation policy frozen across runs [12]. That is the right control for measuring reasoning rather than memorisation, and it also means your older code may not sit in the same distribution as freshly disclosed upstream bugs. The 11.7 billion tokens spread over 960 case-runs, about 12.2 million each [9], so the binding constraint in that harness was the turn cap, not the price list.
Ranked by verification strength, evidence, and original report placement.
Aikido burned 11.7 billion tokens benchmarking the cyber capabilities of 10 AI models, with three attempts each, on 32 fresh off-the-shelf vulnerabilities to rediscover.
The lineup added GLM-5.3, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, Qwen3.8-Max, Kimi K3 and Grok 4.6 to Aikido's earlier known-CVE benchmark.
DeepSeek V4 Pro 0813 finds the most vulnerabilities; pooling three runs reaches 28 of 32.
Three DeepSeek Pro runs cost about $295 and outperform Opus 5, Grok 4.6 or Sol.
Three DeepSeek V4 Flash 0731 runs cost $108 and reach 24 vulnerabilities, matching Grok's best individual pass for less than a quarter of the cost.
DeepSeek Pro found 17 vulnerabilities on its first pass and 28 across three pooled runs.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 29, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Frontier scores at $2/$6: Grok 4.6 ties GPT-5.6 on one evaluator's index for a fifth the output price1 distinct publisher
leadership
Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
build
OpenAI's president says open weights will accelerate the threat. His own cyber model stays gated.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One vendor's numbers, no released artifacts
Every figure in this story — 11.7 billion tokens, 28 of 32, $295, $108, 17 universally solved cases — traces to Aikido's blog and nowhere else. The method is described more candidly than most vendor benchmarks bother with, down to the 30-turn ceiling and the frozen prompts, but no CVE list, per-model table or run log is published, so no one outside the company can re-run a single case. What does verify is the internal arithmetic: ten models across three passes over 32 cases is 960 case-runs, and the stated token bill divides into that at about 12.2 million each.
Inside the issuer's own pipeline
The only deployment on the record is Aikido's: these models were swapped into a bounded copy of the harness behind its AI Code Analysis product, and the company disclosed what the runs cost it. Nobody else has adopted the ranking, replicated the pooling strategy, or priced it in public. The story documents a spend, not an uptake.
Frontier headline, two-case margin
"Open-source models now outperform the public frontier" is doing more work than the data under it. Seventeen of 32 cases fell to every model and one fell to none, leaving 14 that actually rank anything — and inside that band the winning margin is two cases, from three runs, with no variance bounds. Meanwhile the sturdiest finding gets the quietest billing: repetition is real but priced on a curve, roughly $5.80 per finding on the first Pro pass and about $17.90 for each of the 11 that the next two passes added. Aikido leads with the crown and buries the economics.
The harness under test is the product
Aikido sells AI code analysis, and the rig here is a bounded version of that product with contender models dropped into its exploration stage. The conclusion that open models 'harnessed correctly' reach frontier recall is, almost word for word, the pitch for owning the harness. The second tilt is subtler and points the same way: the finding that cheap open models throw the most false leads argues for precisely the triage pipeline Aikido ships downstream of them. None of this makes the runs fake — it does mean the framing and the commercial interest are pointing in one direction, and the only cost figures volunteered are for the two models that flatter the argument.
Specific enough to argue with, thin enough to doubt
This is a benchmark you can actually interrogate: definitions up front, turn limits named, freshness precaution stated, totals that reconcile against the run count. That earns more trust than the usual vendor chart. Holding it down: one interested publisher, three passes per model, nothing released for replication, and a copy of the write-up that cuts off mid-sentence in the traces section, so even the qualitative claims about investigation budgets are only half in view.