Skip to content

Topic

AI coding benchmarks

Tests and evaluation suites that measure how well AI models perform software engineering work such as design, implementation and code review.

Current clusters

build5 publishers

Anthropic prices Sonnet 5.5 at half of Opus 5.5, with Sonnet close behind on most benchmarks and ahead on Terminal-Bench

Anthropic says Claude Sonnet 5.5 lands within about two points of Opus 5.5 on coding and knowledge-work tests at half Opus's per-token price. The 30 percent saving it advertises is against Sonnet 5, so moving agents off Opus depends on per-task token counts nobody has verified independently.

Perspective Coverage

5 publishers
Builder
Builder 43%
Operator
Operator 32%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60