Skip to content

Product1 publisher3 min readPublished

AI vulnerability scanners flag 41% to 99% of safe code in AWS's Deception Benchmark

AWS's Deception Benchmark found AI vulnerability scanners catch up to 95% of real bugs but flag 41% to 99% of safe code. Its samples were built to fool models, so teams still need their own false-alarm count before sizing the triage work.

The Product Desk · Product desk

Illustration accompanying AI vulnerability scanners flag 41% to 99% of safe code in AWS's Deception Benchmark

What happened

  • The Deception Benchmark holds 14,822 code samples written in 16 programming languages and covering more than 70 Common Weakness Enumeration categories.
  • Each sample went through a loop of generating code, testing it against frontier models and hardening it, and samples a model solved easily were removed.
  • Proof-of-exploit prompting cut false positives by 17 to 74 percentage points but missed 7% to 44% of real vulnerabilities.
  • Environment-gated challenges went worse, with models flagging code while ignoring the Kubernetes network policy deployed next to it.
  • No configuration AWS tested kept both false positives and false negatives below 10%.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost Reviewers pay for every false flag in investigation time, and since that load has not been sized anywhere, a team has to count it on its own code before it can staff the queue.
  • decision Turning on proof-of-exploit prompting is a choice about which error to absorb: a quieter review queue in exchange for more real bugs reaching production.
  • constraint A scanner that sees only source code cannot clear findings that depend on deployment settings such as network policies, so those findings keep landing on humans until tools read the environment too.

A scanner flags a server-side request forgery path in a service. A Kubernetes network policy beside that service already blocks the path. AWS built its environment-gated challenges around that kind of setup, running the same code in specific environments to test whether a model notices the policy [10]. Mitch Ashley, vice president and practice lead for software lifecycle engineering at the Futurum Group, said AI scanners that read code without the environment around it cannot tell an exploitable path from a blocked one [16].

After an AI scanner runs, reviewers investigate alarms that turn out to be false, and according to devops.com they spend more time on it than ever [12]. Ashley called false positives verification debt that undermines confidence in AI findings [16]. The same report says it is unclear how much work AI models are creating for these teams [13].

That missing count is why the benchmark's headline range cannot go straight into a staffing plan. The samples that survived AWS's hardening loop are the ones models did not solve easily [4]. In the code-level challenges, a vulnerable version and a safe version differ by a subtle fix, and both look suspicious [11]. The labels were checked with care: reviewers re-examined each one without seeing each other's calls, disagreements went to adjudication, and unresolved cases went to humans [5]. So the range measures code picked because it fools models [4][6]. The report does not name the 12 models [3] or say which one reached 95% detection, and it does not estimate how the rates move on ordinary production code.

Two figures from the range still help with planning. The share of safe code flagged differs by 58 percentage points between the lowest and highest results [1], so model choice alone changes the size of the triage queue. The best detection result still leaves at least 5% of real vulnerabilities unflagged [2]. Neha Rungta, director of applied science at AWS, said the goal is to find the models that produce the fewest false positives [14]. She said that requires models that reason about how a remediation would affect the surrounding environment as well as the existing code [15].

I'd budget reviewer hours before the scanner goes live and treat the false-positive rate as a number the team measures on its own code. The tradeoff is delay: the tool gates nothing until those numbers exist. The measurement is small. A set of vulnerabilities the team has already fixed, scanned in both vulnerable and patched form, copies AWS's paired design [11]. False alarms and misses get counted separately.

From there, two questions place a team on a grid: what a missed bug costs, and how many reviewer hours a week are free to clear false flags. Where misses are expensive and hours are spare, plain scanning without proof-of-exploit prompting fits, with the queue staffed to match. A team with cheap misses and thin hours is the case for proof-of-exploit prompting. It buys a quieter queue at the price of more real bugs slipping through [7]. With cheap misses and spare hours, the mode that measured cheaper on the team's own sample wins. The hard corner is expensive misses and thin hours. No tested configuration covered it [8], so there the scanner sits beside human review as a second reader and does not gate merges.

What to watch

  • AWS publishing per-model results that name the 12 models and pair each detection rate with its false-positive rate.
  • Scanner vendors adding deployment context such as Kubernetes network policies to what their models read before raising a finding.
  • A measured count of triage hours from teams running AI scanners on production code, a figure devops.com says is currently unclear.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories