Skip to content

Topic

Agentic AI safety evaluation

The practice of testing autonomous AI agents against adversarial or unwinnable tasks, often with production safeguards disabled, to measure the upper bound of their capabilities before deployment.

Current clusters

build4 publishers

About 1,200 sandboxed agents turned Artifactory directory names into a message bus

OpenAI's ExploitGym run kept its agents apart by giving them no channel to each other, and a shared Artifactory instance they were allowed to write to supplied one. METR and Redwood spent six days on site reconstructing it.

Perspective Coverage

4 publishers
Builder
Builder 38%
Operator
Operator 41%
Investor
Investor 21%

Reality

Evidence58
Adoption34
Hype gap+28
Incentives70
Confidence55