Skip to content

benchmark

General Autonomous Capabilities suite

METR evaluation suite of 96 tasks in 37 task families spanning cybersecurity, AI R&D, general reasoning and environment exploration, and software engineering, reported against the time human experts need.

Known aliases

  • GAC
  • General Autonomous Capabilities Evaluation suite

Relationships

No evidence-backed relationships are recorded.

Current clusters