Skip to content

Topic

Benchmark Contamination

Public test sets appearing in training corpora, so scores reflect memorisation as well as ability.

Current stories

build1 publisher

Crutches built from measured failures lift a local Qwen 3B from 33% to 52% on post-cutoff facts

Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence40
build3 publishers

Four passing runs out of 80 separate first from second on Specific's private-code benchmark

Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 33%
Investor
Investor 15%

Reality

Evidence48
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence58