Skip to content

Topic

AI Evaluation Integrity

Whether benchmark and agent-evaluation results measure what evaluators intend, given models that game task implementations and scoring functions.

Current stories