Skip to content

benchmark

AutoResearchExam

Benchmark of open-ended machine learning and engineering tasks run over a 24-hour horizon, scoring whether improvements an agent produces still hold on held-out data.

Current clusters