Skip to content

benchmark

HarnessSafe

A benchmark of executable test cases for persistent-state safety risks in LLM agent harnesses, organised around the carriers that hold state between tasks: memory, skills, tools, subagents, summaries and shared artifacts.

Current clusters