Skip to content

benchmark

Fang et al. 15-vulnerability exploit benchmark

Set of fifteen real-world vulnerabilities used to measure whether a language model agent can produce working exploits, run with and without the CVE description supplied.

Current clusters