Skip to content

benchmark

AIME 2025

The 2025 American Invitational Mathematics Examination, used as a competition-maths test for language models.

Current clusters

build1 publisher

Anthropic prices its newer Sonnet a third below Sonnet 4.5

Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.

Publishers:dev.to

Reality

Evidence40
Adoption20
Hype gap+15
Incentives55
Confidence45