security1 publisher
METR clocks the Claude 3.7 Sonnet agent at 50% success on 55-minute expert tasks
METR also had the agent match the median human expert on five AI R&D tasks, using 32 hours of wall clock against the humans' eight, assembled from attempts of two hours or less. The confidence intervals still overlap the other public models it has tested.
Publishers:metr.org
Reality
- Evidence58
- Adoption25
- Hype gap−10
- Incentives40
- Confidence55