Cantina released apex-flash-1, an open-weights vulnerability-research model it says solved 40 of 60 tasks for $2.38, against $74.68 for Claude Opus 5 High. There is no hosted endpoint, so teams download the 321-billion-parameter weights, pay for their own inference and verify the numbers themselves.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence40
Ai2 open-sourced AstaBrief 8B, which writes cited research reports in 51.1 seconds against 178.5 for Asta's Claude-powered mode. Labs that cannot send unpublished research questions to a hosted model can now run a cited-report generator on their own servers.
Perspective Coverage
3 publishers
- Builder
- Builder 55%
- Operator
- Operator 33%
- Investor
- Investor 12%
Reality
- Evidence55
- Adoption18
- Hype gap+22
- Incentives55
- Confidence60
Cantina released Apex Flash-1, an open-weights security model fine-tuned on 50 of its own vulnerability cases. Its only performance evidence is a 60-task evaluation the company ran itself, and a separate build is modified to refuse fewer requests.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence40
Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.
Perspective Coverage
5 publishers
- Builder
- Builder 58%
- Operator
- Operator 33%
- Investor
- Investor 9%
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives70
- Confidence55
Kimi K3's 1.4 terabytes of weights take eight Nvidia GB300s just to sit in memory. Export controls keep those chips away from Moonshot. Modal, Fireworks and Baseten price the hosted result at $3 in and $15 out.
Reality
- Evidence32
- Adoption45
- Hype gap+33
- Incentives74
- Confidence40
The public fight is billed as one about regulatory capture. The operative question is cost incidence, and one lab just showed what an unlegislated security bill looks like.
Reality
- Evidence36
- Adoption24
- Hype gap+32
- Incentives82
- Confidence41
Z.ai says its 743B-parameter GLM-5.3 hits 34.5% on its own code bench using 22% fewer output tokens than GLM-5.2. The weights are still two weeks out.
Reality
- Evidence32
- Adoption18
- Hype gap+38
- Incentives76
- Confidence36