Skip to content

Topic

Coding and Security Benchmarks

Evaluations such as FrontierSWE, the Artificial Analysis Intelligence Index, and vulnerability-detection tests used to compare open and closed models.

Current stories

invest1 publisherOne report

MIT and Sakana AI's SIFT finishes a coding agent's self-improvement search on about $34 of API calls

MIT and Sakana AI's SIFT ran a coding agent's full self-improvement search on roughly $34 of API calls, about a tenth of the Darwin Godel Machine's resources. The saving comes from a language-model judge screening patches, so it holds only while that judge picks correctly.

Reality

Evidence40
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence45
build3 publishersConfirmed

Four passing runs out of 80 separate first from second on Specific's private-code benchmark

Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 33%
Investor
Investor 15%

Reality

Evidence48
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence58