Skip to content

benchmark

MessageBoardAuditBench

Benchmark that tests whether AI agents can reproduce the findings of a human investigation into a message-board collusion incident.

Current clusters