Skip to content

Security1 publisher3 min readPublished

Mandiant found 100 high-severity bugs in two days. Plan for the other side doing the same.

Ten months of production use, 12 assigned CVEs, and a two-day sweep of stolen repositories. Discovery has stopped being the bottleneck; human validation has become one.

The Watch · Security desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Google's Mandiant disclosed the workings of an internal tool called the Agentic Vulnerability Discovery Harness (AVDH), which uses chains of AI agents to hunt for vulnerabilities in source code.
  • Mandiant says AVDH found over 100 verified, high-severity flaws in just two days during a live investigation into stolen corporate repositories.
  • AVDH has been running inside Mandiant for ten months.
  • In that ten-month period AVDH scanned tens of millions of lines of code and produced tens of thousands of findings, according to a blog post published by the Google Threat Intelligence Group.
  • The tool has uncovered dozens of assignable flaws in widely used web extensions and open-source projects, resulting in 12 assigned CVEs, including CVE-2026-13242 and CVE-2026-55803.

Compiled by The WatchSomething wrong?How this is made

Why it matters

Google's Mandiant has published the internals of a tool it calls the Agentic Vulnerability Discovery Harness, a chain of AI agents that reads source code hunting for exploitable flaws, and says the harness produced more than 100 verified, high-severity findings in two days during a live investigation into stolen corporate repositories [1][2]. The number that matters is not 100; it is two days, because the same class of pipeline pointed at the same stolen code by someone else would run just as fast.

The harness has been in use inside Mandiant for ten months, and in that period has scanned tens of millions of lines of code and produced tens of thousands of findings, according to a blog post from the Google Threat Intelligence Group [3][4]. Mandiant researchers Alex Tselevich and Michael Maturi write that the work has yielded 12 assigned CVEs, including CVE-2026-13242 and CVE-2026-55803, with roughly another dozen in active disclosure [5][6]. That is about 24 identifiers either issued or pending [20], against a ten-month average of a little over one assigned CVE per month [21] and a burst rate of at least 50 verified high-severity findings per day during the repository investigation [22]. The disclosure pipeline is orders of magnitude slower than the discovery pipeline.

The architecture is unglamorous and that is its strength. Built on Google's Agent Development Kit, it runs as a sequence of specialised agents, each passing output to the next [7]: a threat modeling stage that maps the codebase and marks what to skip, such as test directories, with a human signing off on the model before anything else runs [8]; entry point discovery across every in-scope file, from web routes to inter-process listeners [9]; context enrichment that gathers the permission checks and sanitizers a reviewer would otherwise chase by hand [10]; hypothesis generation split between access-control problems and dangerous data flows [11]; and validation by several agents deliberately run at high temperature, with a synthesis agent sorting each hypothesis into confirmed, disproven, or rejected [12].

Then a person checks it. Mandiant consultants reproduce the exploit and run proof-of-concept code, and findings that fail are discarded [13]. Tselevich and Maturi explicitly tell defenders building similar harnesses to validate manually [14]. Mandiant says it fought the traditional false-positive problem by having agents challenge each other and test conclusions against rules written by its own consultants, organised by software domain and then by language, framework, and vulnerability type so the knowledge is reusable [15][16]. It also declined to grade itself on public vulnerability datasets, building synthetic vulnerable codebases instead, on the concern that current models may have memorised the public sets during training [17].

That last detail is the most honest thing in the post, and it is worth reading alongside the marketing line. The researchers argue manual review cannot keep pace and that traditional scanners miss too much [18], and that the harness proves defenders can reclaim the advantage against adversarial AI [19]. The first half is a capability claim with evidence behind it. The second is a hope: nothing in the disclosure documents an attacker running such a harness, and the asymmetry cuts the wrong way, since an adversary auditing stolen code does not need reproducible proof-of-concept sign-off before acting.

Watch whether the second dozen CVEs land, and how long assignment takes [6]. Watch whether the organisations whose repositories were stolen can triage at the rate the harness reads. If your source code has ever left your perimeter, treat it as under review by someone with a validation queue shorter than yours.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories