Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.
Reality
- Evidence65
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence62
Flash's off-peak input price is under a quarter of what V4-Pro cost, and on DeepSeek's own table it beats the old Pro checkpoint on Terminal-Bench, but it scores 36.8 on Humanity's Last Exam and no V4.1-Pro has a date.
Reality
- Evidence42
- Adoption52
- Hype gap+30
- Incentives72
- Confidence40
The vendor selling the cheapest model in the comparison reports a 0.7-point quality spread across four frontier models against run-to-run variation of 1.4 to 3.2 points. That leaves price per task, $0.43 against an implied $6.45 for GPT-6 Astra.
Publishers:fireworks.ai
Reality
- Evidence42
- Adoption18
- Hype gap+28
- Incentives88
- Confidence58
A new evaluation gives models explicit rules about what may appear in their reasoning and scores whether they comply while still solving the problem. Compliance rose with model size and fell with more RL training.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives35
- Confidence60
The V4.1-Flash change log puts Terminal Bench 2.1 at 90.6 against V4-Pro's 87.9 and DeepSWE at 74.2 against 62.7. DeepSeek says API prices came down with the release and points to a pricing page for the amounts.
Publishers:api-docs.deepseek.com
Reality
- Evidence45
- Adoption30
- Hype gap+30
- Incentives80
- Confidence55
The 35-billion-parameter Iris-mini and the 397-billion-parameter Iris-pro build on Qwen models and run at 256,000 tokens of context. AllSpark also published the training recipe and scored every benchmark twice, once with context management switched off.
Reality
- Evidence35
- Adoption12
- Hype gap+30
- Incentives70
- Confidence55
Artificial Analysis scores GLM-5.3-Flash 42 at $0.25 a task and Kimi K3 44 at $2.00 a task. At eight-to-one on price, a two-point composite gap settles nothing, and the cost of an hour of human review decides it.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap+12
- Incentives30
- Confidence55
Eleven tasks with pre-computed answer keys, three runs each, seven effort settings. Everything from low upward scored 33 of 33, so the only thing the top rung buys is the number on the launch page.
Reality
- Evidence62
- Adoption45
- Hype gap+15
- Incentives55
- Confidence58
The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.
Perspective Coverage
3 publishers
- Builder
- Builder 27%
- Operator
- Operator 37%
- Investor
- Investor 36%
Reality
- Evidence66
- Adoption32
- Hype gap+61
- Incentives79
- Confidence71
The MIT-licensed 320B model card claims it beats GLM-5.2 at a tenth of the price and approaches Claude Opus 4.8 on coding, but it names no dollar rate, and the comparisons are largely the vendor's own.
Reality
- Evidence34
- Adoption18
- Hype gap+46
- Incentives82
- Confidence61
ARC Prize put GPT-6 Astra at 62.7% against OpenAI's 99.9%, and the third-party composite index has it 0.3 points above the model it replaces, which leaves the 20% safety compute overhead as the clearest number in the launch.
Reality
- Evidence57
- Adoption42
- Hype gap+58
- Incentives74
- Confidence54
Artificial Analysis priced the new model at $10 and $50 per million tokens, and the threefold token cut it measured in the Codex agent harness covers that increase where the roughly 10% cut on its intelligence suite does not.
Reality
- Evidence63
- Adoption
- Insufficient
- Hype gap+16
- Incentives62
- Confidence57
A developer graded 18 expired AI predictions against their own words. The best mark he shows is a C+, and the one he built on took six weeks out of his own roadmap.
Reality
- Evidence30
- Adoption38
- Hype gap+22
- Incentives62
- Confidence32
Z.ai says every gain over GLM-5.2 came from post-training on broader production workflows. Whether that transfers to your stack is not something its private benchmark can tell you.
Perspective Coverage
3 publishers
- Builder
- Builder 45%
- Operator
- Operator 28%
- Investor
- Investor 27%
Reality
- Evidence55
- Adoption30
- Hype gap+25
- Incentives70
- Confidence60
A viral X post said an inference-time text layer put DeepSeek V4 Pro ahead of Fable 5 on every task. The report it points to shows single runs, nine benchmarks, and two losses.
Reality
- Evidence40
- Adoption18
- Hype gap+55
- Incentives62
- Confidence45
The Ultrafast preview runs GPT-5.6 Sol on Cerebras hardware for a hand-picked customer list. That makes capacity allocation, not model choice, the constraint your architecture has to survive.
Perspective Coverage
3 publishers
- Builder
- Builder 42%
- Operator
- Operator 33%
- Investor
- Investor 25%
Reality
- Evidence42
- Adoption24
- Hype gap+32
- Incentives78
- Confidence58
OpenAI's invite-only Ultrafast tier runs the same GPT-5.6 Sol up to 14 times quicker, while Google halves Gemini Flash pricing until December 31. Latency is now its own budget line.
Perspective Coverage
3 publishers
- Builder
- Builder 35%
- Operator
- Operator 33%
- Investor
- Investor 32%
Reality
- Evidence54
- Adoption42
- Hype gap+27
- Incentives74
- Confidence60