Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Cloudflare put frontier models in a harness to mutate attacks its WAF already blocked

Cloudflare ran frontier AI models through 45 scenarios that mutated already-blocked attacks, logging 1,107 attempts and 49 findings after human triage. The findings led to three changes in its Managed Ruleset.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Cloudflare put frontier models in a harness to mutate attacks its WAF already blocked
Generated illustration

What happened

  • A Python harness built and replayed every HTTP request and held scenario state, while one model call proposed each mutation and a second reviewed the response.
  • The models worked blind, with no access to Cloudflare's WAF rules, source code, or internal security signals, so the probing was black-box from their side.
  • Forty-eight of the 49 findings were command injection or server-side request forgery, the two classes the mutation runs kept turning up.
  • Human reviewers were the last gate, checking that each request reached its target, stayed malicious, was clearly unblocked, fell within the WAF's remit, and could be safely reproduced.
  • The three ruleset changes were two new detections, SSRF - Obfuscated Host and SSRF - Restricted Protocol, and an improvement to the existing SSRF - Cloud rule.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability A security team can aim frontier models at its own live defenses and get back a short, triaged list, because the harness runs every request and the models only propose and review.
  • cost The funnel from 1,107 attempts to 49 findings puts human reviewers at the final gate. Review capacity is what bounds how fast fixes ship.
  • precedent As this harness pattern spreads across security tooling, the shippable unit becomes the validated finding. Teams then have to trust the harness that does the validating.

The feedback loop is easiest to see in one SSRF test Cloudflare published. The tester kept rewriting a cloud metadata address, trying decimal, octal, and other forms, and moving where it sat in the request [9]. A decimal version got blocked. On the next attempt the model held the request shape and switched to a trailing-dot form of the address, and the client got a redirect instead of a WAF block [10]. Cloudflare recorded that as a result to investigate, not as proof the attack had worked [11].

Of the 1,107 recorded mutations, 607 reached a triaged result: 558 the firewall blocked and 49 that survived as findings worth remediation work [12]. The other 500 attempts produced nothing a reviewer kept [22]. The 49 findings are about 4.4 percent of everything the harness tried [21].

The 49 findings measure how far a known-blocked attack can be bent before the firewall misses it. The scenarios started from payloads the WAF had already blocked [4], then mutated encoding, placement, and delivery against the responses that came back [5]. To get the same result on another firewall you would need the same inputs: a stock of already-blocked payloads and a harness that keeps request execution out of the model's hands.

The same pattern runs through recent security tooling. The harness constrains execution, keeps state, and validates findings before they count [20]. Cloudflare's own Vulnerability Discovery Harness treats finding a bug and independently confirming it as separate jobs [19]. Google Mandiant's Agentic Vulnerability Discovery Harness hands the work along a line of specialised agents that read the source code, form hypotheses, and check them, and only then passes anything to a human [16]. OpenAI's Codex Security builds a threat model for a repository, hunts for vulnerabilities, and tries to reproduce candidates in an isolated environment before proposing fixes [17]. Google's PageBreak checks whether an AI-generated vulnerability hypothesis is actually exploitable, partly to keep security teams from drowning in plausible but unverified findings [18].

What to watch

  • Whether Cloudflare reports false-positive or regression rates for the two new SSRF detections once they run against live traffic.
  • Whether the harness gets pointed at attack classes beyond SSRF and command injection, which accounted for 48 of 49 findings.
  • Whether other vendors publish comparable attempt-to-finding-to-rule funnels so the 4.4 percent yield can be compared.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories