Google is releasing Gemini 4 Argon, which it says can autonomously find and patch software flaws, only to select partners in its Fairwind program. Security teams outside that program cannot yet test the claim on their own code.
Perspective Coverage
14 publishers
- Builder
- Builder 41%
- Operator
- Operator 33%
- Investor
- Investor 26%
Reality
- Evidence50
- Adoption25
- Hype gap+35
- Incentives70
- Confidence60
Google priced Gemini 4 Argon at $2 and $10 per million input and output tokens, then released it first to trusted cyber defenders in its Fairwind Program. Teams can budget against those rates now but cannot yet measure the token counts they multiply.
Perspective Coverage
10 publishers
- Builder
- Builder 43%
- Operator
- Operator 29%
- Investor
- Investor 28%
Reality
- Evidence62
- Adoption18
- Hype gap+30
- Incentives68
- Confidence66
Google's Gemini 4 Argon matches GPT-6.1 Sol's $2/$10 token price but costs 2.7 times as much per task, according to Artificial Analysis. Argon uses more tokens per job, so buyers still have to compare frontier models by cost per completed task.
Perspective Coverage
4 publishers
- Builder
- Builder 36%
- Operator
- Operator 34%
- Investor
- Investor 30%
Reality
- Evidence68
- Adoption15
- Hype gap+20
- Incentives55
- Confidence65
Xiaomi's public dashboard put the MiMo 2.6 Pro reinforcement-learning run at $1.05 million after about 51 hours, roughly $20,500 an hour. Its restart notes and token count give other teams an all-in reference for pricing their own RL runs.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence55
OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.
Perspective Coverage
5 publishers
- Builder
- Builder 44%
- Operator
- Operator 34%
- Investor
- Investor 22%
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence60
OpenAI's GPT-6.1 Sol halves cached-input pricing to $0.10 per million tokens and leaves standard rates at $2 and $10. Agents that resend long context collect the saving, while other buyers weigh gains shown mostly in OpenAI's own tests.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives65
- Confidence55
Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence50
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
Ox Alpha arrived on OpenRouter free with a million-token window, and OpenRouter says the unnamed provider retains prompts and completions. Coding teams are using it anyway.
Perspective Coverage
3 publishers
- Builder
- Builder 41%
- Operator
- Operator 37%
- Investor
- Investor 22%
Reality
- Evidence62
- Adoption45
- Hype gap+25
- Incentives60
- Confidence55
Ox Alpha is free, undocumented and unclaimed. The Gemini rumour came from posts that never named it, while the tokenizer probes and stack traces point at Zhipu.
Perspective Coverage
6 publishers
- Builder
- Builder 39%
- Operator
- Operator 33%
- Investor
- Investor 28%
Reality
- Evidence55
- Adoption65
- Hype gap+40
- Incentives70
- Confidence55
Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 29%
- Investor
- Investor 18%
Reality
- Evidence58
- Adoption35
- Hype gap+22
- Incentives72
- Confidence62
The Copilot research preview picks a single, cascade, or critique workflow per request, and you pay standard Copilot rates for every token in every leg. So the router has to save more expensive inference than the extra calls cost.
Perspective Coverage
3 publishers
- Builder
- Builder 58%
- Operator
- Operator 28%
- Investor
- Investor 14%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence60
Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.
Reality
- Evidence65
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence62
OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.
Perspective Coverage
8 publishers
- Builder
- Builder 36%
- Operator
- Operator 42%
- Investor
- Investor 22%
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence60
OpenAI's own benchmarks put GPT-6 Sol and Luna at half the list price and a fraction of a rival's cost per finished task. For defenders, the thing getting cheaper is autonomous tool calls into business systems.
Reality
- Evidence32
- Adoption38
- Hype gap+34
- Incentives80
- Confidence52
Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.
Reality
- Evidence42
- Adoption55
- Hype gap+25
- Incentives78
- Confidence52
Two agentic post-training jobs, a 1.02-trillion-parameter model and a 309-billion cousin, ran with step time, reward and infrastructure error rates on the open web. Both stopped near step 30, and both models are still unreleased.
Publishers:rajeshparikh.substack.com
Reality
- Evidence34
- Adoption14
- Hype gap+31
- Incentives62
- Confidence41
xAI built Grok 4.7 on a larger base model with a longer reinforcement-learning run and kept the API at Grok 4.6's rates. The open question for buyers is how many tokens the longer runs burn.
Reality
- Evidence34
- Adoption22
- Hype gap+28
- Incentives82
- Confidence46
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
Publishers:forkast.news · openai.com Reality
- Evidence55
- Adoption32
- Hype gap+30
- Incentives82
- Confidence62
Artificial Analysis scores the new Xiaomi model first among open weights and twelve points behind Claude Opus 5.5, and the cheaper Flash tier is the one an operator should put in front of a real queue.
Reality
- Evidence55
- Adoption30
- Hype gap+15
- Incentives72
- Confidence55
Earlier coverage
- Artificial Analysis retries a provider safety error ten times before scoring the attempt zero
Build · September 19, 2026 · 1 publisher
- Real-SWE licenses private production codebases to score coding agents on real business tasks
Build · September 18, 2026 · 1 publisher
- Fireworks' own DeepSWE numbers put four coding models inside the noise band
Product · September 17, 2026 · 1 publisher
- DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores
Invest · September 14, 2026 · 1 publisher
- DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters
Leadership · September 12, 2026 · 1 publisher
- DeepSeek reroutes V4-Pro API traffic to a smaller model on September 14
Product · September 11, 2026 · 1 publisher
- Cognition put a cost penalty inside SWE-2's reinforcement-learning objective
Build · September 10, 2026 · 1 publisher
- GitHub's HydraFusion turns model selection into a routing decision Copilot makes for you
Product · September 8, 2026 · 1 publisher
- GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor
Leadership · September 5, 2026 · 1 publisher
- Astra's 99.9% holds up only on the harness OpenAI ran itself
Invest · September 4, 2026 · 1 publisher
- Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test
Build · August 31, 2026 · 2 publishers
- Harness choice moved token use 83-fold with the model held constant
Build · August 27, 2026 · 1 publisher
- GLM-5.3 changed nothing but the training environments. That is the whole test.
Build · August 19, 2026 · 3 publishers
- Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design
Build · August 19, 2026 · 2 publishers
- Unsloth's 10% quant claim is really about which machines can run a 27B model
Build · August 19, 2026 · 1 publisher
- A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly
Build · August 17, 2026 · 1 publisher
- Three frontier launches in a day, all pitched on price. Open weights set the ceiling.
Build · August 14, 2026 · 4 publishers