Skip to content

benchmark

DeepSWE

Software-engineering benchmark on which Grok 4.6 is reported as trailing the field.

Known aliases

  • datacurve/deep-swe
  • Deep SWE
  • DeepSWE 1.1
  • DeepSWE v1.1

Relationships

No evidence-backed relationships are recorded.

Current stories

product14 publishers

Google limits Gemini 4 Argon to select partners in its Fairwind security program

Google is releasing Gemini 4 Argon, which it says can autonomously find and patch software flaws, only to select partners in its Fairwind program. Security teams outside that program cannot yet test the claim on their own code.

Perspective Coverage

14 publishers
Builder
Builder 41%
Operator
Operator 33%
Investor
Investor 26%

Reality

Evidence50
Adoption25
Hype gap+35
Incentives70
Confidence60
build10 publishers

Google publishes Gemini 4 Argon's token prices before most teams can call the model

Google priced Gemini 4 Argon at $2 and $10 per million input and output tokens, then released it first to trusted cyber defenders in its Fairwind Program. Teams can budget against those rates now but cannot yet measure the token counts they multiply.

Perspective Coverage

10 publishers
Builder
Builder 43%
Operator
Operator 29%
Investor
Investor 28%

Reality

Evidence62
Adoption18
Hype gap+30
Incentives68
Confidence66
science4 publishers

Gemini 4 Argon costs 2.7 times as much per task as GPT-6.1 Sol at the same token price

Google's Gemini 4 Argon matches GPT-6.1 Sol's $2/$10 token price but costs 2.7 times as much per task, according to Artificial Analysis. Argon uses more tokens per job, so buyers still have to compare frontier models by cost per completed task.

Perspective Coverage

4 publishers
Builder
Builder 36%
Operator
Operator 34%
Investor
Investor 30%

Reality

Evidence68
Adoption15
Hype gap+20
Incentives55
Confidence65
product5 publishers

OpenAI's 10-cent GPT-6.1 Sol price covers only cached input

OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.

Perspective Coverage

5 publishers
Builder
Builder 44%
Operator
Operator 34%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
build1 publisher

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
build4 publishers

DeepSeek open-sources the harness, then raises the price of the model

Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.

Perspective Coverage

4 publishers
Builder
Builder 51%
Operator
Operator 31%
Investor
Investor 18%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence58
build6 publishers

A million-token stranger on OpenRouter, and 30 of 30 tokenizer matches with GLM-5.3

Ox Alpha is free, undocumented and unclaimed. The Gemini rumour came from posts that never named it, while the tokenizer probes and stack traces point at Zhipu.

Perspective Coverage

6 publishers
Builder
Builder 39%
Operator
Operator 33%
Investor
Investor 28%

Reality

Evidence55
Adoption65
Hype gap+40
Incentives70
Confidence55
build8 publishers

Gemini 3.8 Flash's introductory price doubles on December 31, 2026

Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 29%
Investor
Investor 18%

Reality

Evidence58
Adoption35
Hype gap+22
Incentives72
Confidence62
build2 publishers

GitHub bills HydraFusion by every model leg its router decides to call

The Copilot research preview picks a single, cascade, or critique workflow per request, and you pay standard Copilot rates for every token in every leg. So the router has to save more expensive inference than the extra calls cost.

Perspective Coverage

3 publishers
Builder
Builder 58%
Operator
Operator 28%
Investor
Investor 14%

Reality

Evidence55
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence60
product8 publishers

OpenAI halves the API price of Sol and Luna against GPT-5.6's promotional rates

OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.

Perspective Coverage

8 publishers
Builder
Builder 36%
Operator
Operator 42%
Investor
Investor 22%

Reality

Evidence40
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
product1 publisher

OpenAI cuts prices on new GPT-6 Sol and Luna models

Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.

Publishers:thenextweb.com

Reality

Evidence42
Adoption55
Hype gap+25
Incentives78
Confidence52

Earlier coverage

  1. Artificial Analysis retries a provider safety error ten times before scoring the attempt zero

    Build · September 19, 2026 · 1 publisher

  2. Real-SWE licenses private production codebases to score coding agents on real business tasks

    Build · September 18, 2026 · 1 publisher

  3. Fireworks' own DeepSWE numbers put four coding models inside the noise band

    Product · September 17, 2026 · 1 publisher

  4. DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores

    Invest · September 14, 2026 · 1 publisher

  5. DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters

    Leadership · September 12, 2026 · 1 publisher

  6. DeepSeek reroutes V4-Pro API traffic to a smaller model on September 14

    Product · September 11, 2026 · 1 publisher

  7. Cognition put a cost penalty inside SWE-2's reinforcement-learning objective

    Build · September 10, 2026 · 1 publisher

  8. GitHub's HydraFusion turns model selection into a routing decision Copilot makes for you

    Product · September 8, 2026 · 1 publisher

  9. GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor

    Leadership · September 5, 2026 · 1 publisher

  10. Astra's 99.9% holds up only on the harness OpenAI ran itself

    Invest · September 4, 2026 · 1 publisher

  11. Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test

    Build · August 31, 2026 · 2 publishers

  12. Harness choice moved token use 83-fold with the model held constant

    Build · August 27, 2026 · 1 publisher

  13. GLM-5.3 changed nothing but the training environments. That is the whole test.

    Build · August 19, 2026 · 3 publishers

  14. Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design

    Build · August 19, 2026 · 2 publishers

  15. Unsloth's 10% quant claim is really about which machines can run a 27B model

    Build · August 19, 2026 · 1 publisher

  16. A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly

    Build · August 17, 2026 · 1 publisher

  17. Three frontier launches in a day, all pitched on price. Open weights set the ceiling.

    Build · August 14, 2026 · 4 publishers