Skip to content

Benchmark

FrontierCode 1.1

Coding benchmark cited in an early-access partner report where Opus 5 approaches Fable-level performance at half the cost.

Known aliases

  • FrontierCode
  • FrontierCode 1.1
  • FrontierCode 1.1 Main

Current stories

buildConfirmed8 publishers

Gemini 3.7 Flash's real pitch is fewer dead agent runs, and it is half price until December

Google's new workhorse Flash is a model-string swap for anyone on AI Gateway, discounted through 31 December 2026. The claim worth testing is reduced tool-calling loop failures, not benchmark deltas.

Perspective Coverage

8 publishers
Builder
Builder 54%
Operator
Operator 24%
Investor
Investor 22%

Reality

Evidence55
Adoption45
Hype gap+25
Incentives70
Confidence60
buildConfirmed2 publishers

Cognition reports 4.8x token throughput on Nvidia Vera Rubin in its own coding-agent test

Cognition says its SWE-2 model produced up to 4.8 times more total token throughput on Nvidia's Vera Rubin NVL72 than on GB200, in a test on CoreWeave. The company-run result points to more agent capacity per system, but whether long coding tasks get cheaper depends on what the new systems cost per hour.

Reality

Evidence40
Adoption30
Hype gap+35
Incentives85
Confidence60