Skip to content

Topic

Simulation Evaluation and Accuracy Claims

How agreement with human panels is measured and reported, and why a wide accuracy band hides very different error rates.

Current stories

buildOne report1 publisher

Cantina's open-weights exploit model ran a 60-task security eval for $2.38

Cantina released apex-flash-1, an open-weights vulnerability-research model it says solved 40 of 60 tasks for $2.38, against $74.68 for Claude Opus 5 High. There is no hosted endpoint, so teams download the 321-billion-parameter weights, pay for their own inference and verify the numbers themselves.

Publishers:runtimewire.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence40