Skip to content

benchmark

Fang et al. (2024)

Peer-reviewed 2024 study measuring how often LLM-driven agents can exploit real one-day vulnerabilities without human assistance.

Current clusters