Skip to content

benchmark

SWE-bench Pro

Software-engineering benchmark reported by Qwen at 61.7, above Opus 4.6 Max's 53.4 in the same table.

Current stories

build3 publishers

Four passing runs out of 80 separate first from second on Specific's private-code benchmark

Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 33%
Investor
Investor 15%

Reality

Evidence48
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence58
build1 publisher

Anthropic prices its newer Sonnet a third below Sonnet 4.5

Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.

Publishers:dev.to

Reality

Evidence40
Adoption20
Hype gap+15
Incentives55
Confidence45
build5 publishers

Alibaba ships the Qwen4 architecture as open weights before the flagship exists

Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.

Perspective Coverage

5 publishers
Builder
Builder 52%
Operator
Operator 28%
Investor
Investor 20%

Reality

Evidence58
Adoption52
Hype gap+32
Incentives76
Confidence71
build1 publisher

Cached prefixes push 91% of a ReAct agent's LLM time into decode

An arXiv tracing study of Claude Code agents on Gemma and Qwen measured prefix-cache hit rates between 84.6 and 99.5 percent, which moves the serving bottleneck to how long you can keep KV blocks resident between tool calls.

Publishers:arxiv.org

Reality

Evidence58
Adoption25
Hype gap+12
Incentives
Insufficient
Confidence55