Published Security3 min read
Most teams cannot triage 500 AI findings. That is a staffing number, not a tooling one
A survey of 158 practitioners found only 20.3% have a workflow for high-volume AI pentest output. Verification capacity, not tool selection, is the binding constraint.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- A survey asked 158 practitioners whether their teams could triage more than 500 AI-generated vulnerability candidates from a single engagement.
- Only 20.3% of the 158 respondents said they had a workflow in place to handle more than 500 AI-generated vulnerability candidates.
- 38.6% of respondents said that volume of findings would strain the team.
- 29.7% of respondents said that volume of findings was unmanageable.
- 68.3% of respondents said the volume would either strain the team or be unmanageable.
Compiled by The WatchSomething wrong?How this is made
Why it matters
A survey of 158 security practitioners, reported by Security Affairs, asked a blunt capacity question: could your team triage more than 500 AI-generated vulnerability candidates from a single engagement? Only 20.3 percent said they had a workflow in place, while 38.6 percent said that volume would strain the team and 29.7 percent called it unmanageable [1][2][3][4]. Taken together, 68.3 percent of respondents expect pain or failure at a volume that current tools produce routinely [5], and 79.7 percent have no workflow for it at all [6].
The article's framing is that discovery has stopped being the bottleneck and verification has become one. It calls the resulting backlog validation debt: unverified findings piling up because discovery scaled and verification did not [17]. The mechanics are unglamorous. Someone has to reproduce the finding, establish whether it is exploitable, and hand engineering enough evidence to act [21]. Severity ratings do not substitute for that, because a high score on an asset that cannot hurt the business tells you very little [22].
The numbers that matter here are hours, not accuracy percentages. Security Affairs cites one respondent who spent two days validating 300 findings from an AI tool, of which 250 were duplicates, non-exploitable issues, or references to vulnerabilities that did not exist [8]. That is an 83.3 percent junk rate, with 50 findings left standing after two days of work [9], or roughly 3.2 minutes of analyst time per candidate [10]. The article's own estimate is more generous and still ugly: at five minutes each, validating 1,000 findings runs past 80 hours [11], about 83 hours of labour [12]. Apply the same rate to the survey's 500-candidate threshold and you get roughly 42 hours, one analyst's full week, spent deciding what is real [23].
That is why tool price is the wrong line item to argue about. Security Affairs poses the arithmetic directly: if automation removes 10 hours from discovery but adds 15 hours of validation and triage, the only thing acquired is validation debt [19]. And 81.7 percent of practitioners using AI tools reported findings that needed significant manual validation at least sometimes [7]. The output makes this harder to catch early, because polished descriptions, severity ratings, attack narratives and remediation advice read as authoritative without being proof [18].
There is a maturity split in the data worth noting. Among teams running fewer than five tests per month, only 4 percent reported a formal workflow for high volumes of AI-generated findings and 55 percent said the volume would be unmanageable [13]. Those low-frequency testers are about a fifth as likely to have a workflow as the survey population overall [14]. The article's reading is that frequent testers cope better because they were forced to build triage discipline [15], though it also cautions that the sample thins out at the highest testing frequencies, so treat it as a trend [16]. Security Affairs also notes the industry is busy measuring how fast AI finds vulnerabilities, naming Anthropic, while far fewer people cost out who checks the work [20].
What to watch: whether vendors start publishing validated-finding rates rather than raw discovery counts, and whether buyers ask for a duplicate and non-exploitable rate before signing. The internal signal is simpler. If nobody owns the triage queue by name, the capacity question has already been answered.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A survey asked 158 practitioners whether their teams could triage more than 500 AI-generated vulnerability candidates from a single engagement.
- [2]
Only 20.3% of the 158 respondents said they had a workflow in place to handle more than 500 AI-generated vulnerability candidates.
- [3]
38.6% of respondents said that volume of findings would strain the team.
- [4]
29.7% of respondents said that volume of findings was unmanageable.
- [7]
81.7% of practitioners using AI tools discovered findings that needed significant manual validation at least sometimes.
- [8]
One survey respondent spent two days validating 300 findings from an AI tool; 250 of them were duplicates, non-exploitable issues, or references to vulnerabilities that did not exist.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- securityaffairs.comPierluigi PaganiniAug 11The inconvenient truth about AI pentesting: someone has to check all the work
Additional citations
- Security Affairs



