Skip to content

benchmark

SWE-bench Verified

A human-validated subset of SWE-bench that evaluates AI coding agents on real GitHub issues from open-source Python repositories.

Known aliases

  • SWE-bench
  • SWE-bench verified
  • SWE-bench-verified

Relationships

No evidence-backed relationships are recorded.

Current stories

build3 publishers

Kimi K3 on its cheapest host undercuts Fireworks' Ember-1 despite a 23% cut in reasoning tokens

Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 30%
Investor
Investor 18%

Reality

Evidence55
Adoption30
Hype gap+25
Incentives70
Confidence58
build1 publisher

Asking GPT-5.6 Luna to name an amphibian flags benchmark transcripts with black-box access

GPT-5.6 Luna says "frog" 70-95% of the time when asked for an amphibian after capability benchmarks, against 12-38% after real use, a LessWrong post reports. Anyone with black-box access can run the check, though its authors cannot yet say whether it detects evaluation awareness or lexical cues.

Publishers:lesswrong.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence40
build1 publisher

Anthropic prices its newer Sonnet a third below Sonnet 4.5

Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.

Publishers:dev.to

Reality

Evidence40
Adoption20
Hype gap+15
Incentives55
Confidence45

Earlier coverage

  1. Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design

    Build · August 19, 2026 · 2 publishers

  2. Ornith-1.0's benchmarks are fine. Ollama can't parse its tool calls.

    Build · August 18, 2026 · 1 publisher

  3. The AI bill nobody reconciles: cost per finished task, not per million tokens

    Leadership · August 18, 2026 · 1 publisher

  4. Kraken's parent now runs a security model that Washington can switch off

    Invest · August 17, 2026 · 1 publisher

  5. NIST says AI benchmarks are now an attack surface, not just a measuring stick

    Science · August 16, 2026 · 1 publisher

  6. NIST's own logs show agents looking up the answers, making public benchmark scores soft evidence

    Science · August 16, 2026 · 1 publisher