Security1 distinct publisher3 min readPublished
Agent accounts clustered in 17 of the Top 25 finishers at the 2026 benchmark while producing a fraction of the solves. The talent-substitution pitch needs the opposite pattern.
The Watch · Security desk
Compiled by The WatchSomething wrong?How this is made
Divide the two percentages and the amplification reading gets sharper than the headline count. Agent accounts were 2.7% of registrations and produced 4.2% of submitted flags, so each agent seat carried about 1.6 times its share of solves [3][1]. Only 46 of the 93 were active at all, and if the dormant ones submitted nothing, the working agents ran at roughly three times their registration weight [2][2]. That is a genuine productivity figure. It is also a figure about accounts, not about practitioners, and the accounts sat with teams that were already finishing near the top.
The distribution is where a substitution case would have to be made, and it is not there. Seventeen of the Top 25 finishers carried an agent; 33 of the Top 100 did [4][5]. Subtract, and among teams placing 26th through 100th just 16 of 75 had one, about 21% against 68% inside the Top 25 [3]. Widen the frame and it gets worse: 54 teams fielded agents, 33 of those landed in the Top 100, which leaves 21 agent-carrying teams, close to two in five, outside it [2][5][4]. Bringing the tool did not place anyone.
The head-to-head from a year earlier points the same direction. At the NeuroGrid CTF in November 2025, Hack The Box put AI-augmented and human-only teams on one challenge set; across the whole field the AI side solved at 3.2 times the rate, and among the Top 5% that multiple fell to 1.69 [11]. The edge among the best is roughly half the edge across the field [5]. A tool that replaces scarce senior judgment would show the reverse gradient, doing most for the teams that know least. This one does most for the middle and least for the people who are hardest to hire.
Speed is the one axis where the agents clearly won: the best AI-augmented teams worked three to four times faster [12]. They did not finish. The only team to clear all 36 challenges was human, and the strongest AI team stopped at 32, four short [13][7].
The longitudinal numbers are being quoted more confidently elsewhere than Hack The Box quotes them itself. Median solve time was 26 hours into the event two years ago and 13.8 hours in 2026, a drop of 12.2 hours or about 47% [8][6]. Full completions went from two teams in 2024 to three in 2025 to fifteen in 2026, a factor of 7.5 [10][8]. Hack The Box's own caveat is that board size, challenge mix, who turned up, how they prepared and what tooling they had all changed over that period and the data cannot separate them [9]. The company also says plainly that proximity to the top is not causation [7]. Coming from a vendor with an agent story available to tell, that is the more useful part of the release.
CEO Haris Pylarinos frames it as AI appearing alongside the strongest practitioners rather than instead of them, and argues that human judgment, validation and hands-on skill become more important as agents improve [6]. The telemetry supports the first half directly. The second half has a cost attached that the percentages hint at: 4.2% of an event's flags arriving through 46 machine accounts is also a review queue, and the people qualified to work it are the 17 teams in the Top 25 who were already the constraint [2][3][4].
Ranked by verification strength, evidence, and original report placement.
Agent accounts held 2.7% of registered accounts and produced 4.2% of submitted flags and 4.6% of awarded points.
Agents showed up in 17 of the Top 25 finishers in the 2026 Global Cyber Skills Benchmark.
Thirty-three of the Top 100 teams had an agent account.
Two teams cleared every challenge in 2024, three did in 2025 and fifteen did in 2026; board design also changed, so the count tracks breadth, coordination and endurance rather than event difficulty.
Hack The Box ran the 2026 Global Cyber Skills Benchmark under the name Project Nightfall.
The event counted 93 designated agent accounts across 54 teams, 46 of them active.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific but single-source and unaudited
The figures are unusually granular — account, flag and point shares, leaderboard placement splits, two solve-rate multiples from a same-challenge-set comparison — and the publisher preserves the vendor's own confounder caveats, which raises credibility above a pure press release. But every number originates with Hack The Box's instrumentation of its own competitions, there is no independent verification, no absolute denominators, no definition or enforcement detail for what counted as an 'agent account', and no disclosure of models or agent frameworks. The strongest internal check is that the vendor publishes results that partly cut against a simple AI-superiority story (human-only full clear, advantage compressing at the top).
Thin overall, concentrated at the elite tier
Agent use is real and disclosed with numbers rather than asserted, but it is small in absolute terms: 2.7% of registered accounts, 46 active agent accounts, 4.2% of flags. Penetration is meaningful only in a narrow band — 68% of the Top 25 and 33 of the Top 100 teams — while about 39% of agent-carrying teams finished outside the Top 100 and the mid-field adopted at roughly 21%. This is early, uneven, competition-context adoption with no evidence supplied about production security operations.
Substitution narrative runs ahead of the numbers
Positive because the prevailing claim these figures are usually recruited to support — that agents substitute for scarce security talent — is not what the data shows: agents produced a low-single-digit share of output, concentrated with the practitioners who needed help least, lost the completeness contest to a human-only team, and saw their solve-rate edge compress from 3.2x to 1.69x at the top of the field. The gap is not large because both the vendor and the publisher hedge rather than oversell: the correlation-not-causation caveat and the board-design confounders are stated plainly, and the two most quotable trend lines (solve time halving, 7.5x more full clears) are explicitly declared causally unattributable in the same breath.
Vendor-generated data with an aligned conclusion and a gated asset
All data is produced, curated and released by Hack The Box, whose commercial business is human cyber-skills training and assessment; the conclusion the release advances — that as agents improve, human judgment and hands-on skill matter more — is the conclusion most favorable to that business, delivered as a CEO quote. The article itself terminates in a 'Download: The high-performance team playbook' call to action, marking a lead-generation pathway. Partially offsetting: the vendor published findings that constrain its own AI narrative and volunteered confounders, which a purely promotional release would omit.
Moderate: internally consistent, externally unchecked
Confidence is mid-range. The arithmetic derivations follow cleanly from stated figures, the two events point the same direction, and the caveats are on the record, so the internal story is coherent. But a single publisher relaying a single vendor's unaudited telemetry, with no primary methodology document, no agent-designation criteria and no corroboration, caps how far these numbers can be relied on for decisions about staffing or agent capability.
security
GitLab 19.3 puts agent runtime, inference models and secrets under one permission model1 distinct publisher
security
An agent guard that runs on your laptop, and cannot tell you whether anyone keeps it on1 distinct publisher
security
OpenAI's Computer History writes a plaintext log of the workday. Decide before staff opt in.1 distinct publisher
security
The customer is genuine and the payment is authorized: 55% of banks say scams dominate fraud1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026