build1 publisher
Shihipar's runs show Claude Code's effort level mostly buys self-testing
Thariq Shihipar's runs show Claude Code at max effort cut missed edge cases from 59 to 24, yet misread unclear tasks 47 times against 25 at low. At about triple the tokens per attempt, a high default pays for self-checking that only some kinds of work turn into passes.
Publishers:dev.to
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence35