product1 publisher
Splitting a data room across sub-agents beats a standard tool-loop harness on Harvey's benchmark
Harvey and Baseten report rubric pass rates rising from 23.3% to 62.4% across seven models once the data room sits in a Python REPL and sub-agents do the reading, while Claude Code with Opus-5 manages 24.6%.
Publishers:harvey.ai
Reality
- Evidence47
- Adoption17
- Hype gap+18
- Incentives84
- Confidence57