Skip to content

Topic

AI Vulnerability Discovery and Exploitation

Model capability at finding vulnerabilities in source code and chaining exploitation stages, measured by CyberGym, ExploitBench and ExploitGym.

Current stories

security2 publishers

Epic pauses most development for about six weeks over MyChart flaws an AI model found

Epic has stopped most product development for roughly six weeks while it fixes MyChart flaws that an AI model uncovered, according to TechCrunch. No list of affected portal setups is public yet, so hospitals cannot check theirs against one.

Perspective Coverage

3 publishers
Builder
Builder 30%
Operator
Operator 57%
Investor
Investor 13%

Reality

Evidence55
Adoption25
Hype gap+15
Incentives40
Confidence60
invest2 publishers

Palo Alto Networks built a service on the Anthropic model that found 75 flaws in its own products

Palo Alto Networks pointed Anthropic's unreleased Mythos at its own systems and found 75 vulnerabilities in a month, against a usual rate below five. The defense business it built on that result depends on Anthropic's model and on customers choosing a security vendor over the lab.

Reality

Evidence45
Adoption30
Hype gap+35
Incentives80
Confidence50
security1 publisher

Chainguard discloses 14 Java bugs that were fixed upstream but never got a CVE

Chainguard disclosed 14 Java vulnerabilities that were fixed upstream but never assigned a CVE, one rated critical and one high. Teams still on the affected versions got no scanner alert, some for years, because nobody announced the fixes when they landed at HEAD.

Publishers:chainguard.dev

Reality

Evidence40
Adoption25
Hype gap+10
Incentives75
Confidence45
build5 publishers

GLM-5.3 keeps GLM-5.2's base model and claims 50% more on coding: plan for shorter eval cycles

Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.

Perspective Coverage

5 publishers
Builder
Builder 58%
Operator
Operator 33%
Investor
Investor 9%

Reality

Evidence40
Adoption30
Hype gap+35
Incentives70
Confidence55

Earlier coverage

  1. AI bug-hunting models pushed Microsoft's four-month patch total to 4.5 times its old baseline

    Product · September 8, 2026 · 1 publisher

  2. Sysdig credits Anthropic's Mythos preview with 181 working Firefox exploits

    Security · September 6, 2026 · 1 publisher

  3. Arista tells network teams to staff for months of batched EOS and VeloCloud advisories

    Security · September 6, 2026 · 1 publisher

  4. August's 398-CVE Patch Tuesday moves the bottleneck to the test bench

    Security · September 4, 2026 · 1 publisher

  5. Maintainers shipped 97 fixes against the 23,019 bugs Claude Mythos flagged

    Security · September 3, 2026 · 1 publisher

  6. A bug drought would push backdoor mandates back onto product roadmaps

    Product · August 31, 2026 · 1 publisher

  7. The exploit was the toll gate: Gartner's top emerging risk moved five places in one quarter

    Security · August 26, 2026 · 1 publisher

  8. Mandiant found 100 high-severity bugs in two days. Plan for the other side doing the same.

    Security · August 19, 2026 · 1 publisher

  9. Cheap bug-hunting arrives: GLM 5.3 puts near-frontier vulnerability discovery on your own hardware

    Product · August 18, 2026 · 1 publisher

  10. Anthropic found 10,000 critical bugs. The bottleneck is now the person reading the report

    Build · August 17, 2026 · 1 publisher

  11. Z.ai's GLM-5.3 beats Claude on CyberGym, then hands out the weights

    Product · August 15, 2026 · 1 publisher

  12. GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead

    Invest · August 14, 2026 · 1 publisher