Skip to content

Topic

AI Benchmarking and Verification

Third-party versus vendor-run measurement of throughput, wall-clock task time and quality retention.

Current stories

invest4 publishers

Sonnet 5.5's extra tokens shrink its half-price edge over Opus to roughly a fifth at max effort

Anthropic's Claude Sonnet 5.5 beats Opus 5.5 at coding for half the per-token price, on its own tests and on Artificial Analysis's. At max effort it writes 60% more tokens per task, so moving coding work down a tier saves nearer a fifth than a half.

Perspective Coverage

4 publishers
Builder
Builder 39%
Operator
Operator 36%
Investor
Investor 25%

Reality

Evidence60
Adoption30
Hype gap+25
Incentives65
Confidence58
product1 publisher

Vals sells a private AI exam to the labs it grades

Vals keeps its test materials private and charges the model developers it scores, and Andreessen Horowitz has now put $40 million behind that arrangement. Buyers reading the numbers cannot run the test themselves.

Publishers:techcrunch.com

Reality

Evidence33
Adoption35
Hype gap+34
Incentives82
Confidence46