GMI Cloud raised $663 million: a $223 million Series B with Nvidia in the round and a $440 million credit facility from ChinaTrust Commercial Bank. Teams that need GPUs in Taiwan, Thailand or Malaysia get a funded local supplier whose expansion rests mostly on borrowed money.
Perspective Coverage
3 publishers
- Builder
- Builder 20%
- Operator
- Operator 37%
- Investor
- Investor 43%
Reality
- Evidence55
- Adoption50
- Hype gap+30
- Incentives65
- Confidence60
g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence45
GMI Cloud raised $668 million, two-thirds of it through a $445 million credit facility led by CTBC and the rest as a $223 million Series B. Repaying debt on that scale depends on the more than $600 million in contracted annual revenue GMI reports and on how quickly it puts GPUs into production to serve it.
Reality
- Evidence35
- Adoption50
- Hype gap+20
- Incentives65
- Confidence40
Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 30%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption30
- Hype gap+25
- Incentives70
- Confidence58
The inference challenger now sells Nvidia clusters and takes Nvidia money. Buyers who priced in a non-Nvidia alternative should re-read their renewal terms.
Perspective Coverage
5 publishers
- Builder
- Builder 19%
- Operator
- Operator 33%
- Investor
- Investor 48%
Reality
- Evidence68
- Adoption40
- Hype gap+25
- Incentives72
- Confidence65
Tenet is post-trained on Moonshot's Kimi K3, converting per-call payments to OpenAI, Anthropic and Google into a fixed training bill. The counterparty risk moved rather than disappeared.
Publishers:harvey.ai · thenextweb.com Reality
- Evidence55
- Adoption15
- Hype gap+30
- Incentives60
- Confidence55
Stratechery argues that slowing model improvement would shrink overhangs the labs' own speed created. The same piece says the safety motive behind Anthropic's pacing position is sincere, so both readings stand.
Reality
- Evidence35
- Adoption25
- Hype gap+25
- Incentives55
- Confidence45
Kimi K3's 1.4 terabytes of weights take eight Nvidia GB300s just to sit in memory. Export controls keep those chips away from Moonshot. Modal, Fireworks and Baseten price the hosted result at $3 in and $15 out.
Reality
- Evidence32
- Adoption45
- Hype gap+33
- Incentives74
- Confidence40
The vendor selling the cheapest model in the comparison reports a 0.7-point quality spread across four frontier models against run-to-run variation of 1.4 to 3.2 points. That leaves price per task, $0.43 against an implied $6.45 for GPT-6 Astra.
Publishers:fireworks.ai
Reality
- Evidence42
- Adoption18
- Hype gap+28
- Incentives88
- Confidence58
Apple's Kids Category rule and COPPA's definition of a child's voice both point a preschool voice product at local processing, and Whisper's error on children's speech falls by nearly a factor of three from tiny to large-v3.
Reality
- Evidence47
- Adoption38
- Hype gap+9
- Incentives71
- Confidence54
One endpoint fronts more than 300 models, but the company running the GPUs picks the inference engine and the quantization. A dev.to writeup says the quality gap that follows turns up in the response body, while the status code still reads success.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Arize and Fireworks priced ten models on what a completed command-line task costs, across 2,400 runs. The winner on that metric is an open model with the worst pass rate in the study and the thinnest coverage.
Publishers:arize.com
Reality
- Evidence62
- Adoption18
- Hype gap+14
- Incentives75
- Confidence55
Groq kept its data centers and its customer contracts while most of its engineers became Nvidians inside eight days, and the three customers who spoke on the record disagree about what that cost them.
Reality
- Evidence54
- Adoption57
- Hype gap+14
- Incentives71
- Confidence55
Open-weight models take about 60 percent of the tokens on US-originating OpenRouter requests, and most of those weights come from Chinese labs, so the marketplace is now selling residency of the inference itself.
Reality
- Evidence55
- Adoption35
- Hype gap+20
- Incentives65
- Confidence58
The operator boards copied for twenty years spent 1993 to 2003 in product and engineering jobs. SaaStr counts 82 of the 100 fastest-growing AI-native startups with a technical CEO, against 49% of the 2013 unicorn class.
Reality
- Evidence45
- Adoption58
- Hype gap+28
- Incentives55
- Confidence40
Baseten's case for owning the agent runtime is that latency and idle compute accumulate at every boundary an agent crosses between the model call and its sandbox, on numbers Blaxel reported itself.
Reality
- Evidence42
- Adoption30
- Hype gap+26
- Incentives78
- Confidence52
The $15.5 billion price works out near 39 times annual recurring revenue above $400 million, and it rests on a margin claim Harvey has not quantified: that tuning Kimi K3 beats paying OpenAI per call.
Reality
- Evidence32
- Adoption61
- Hype gap+30
- Incentives74
- Confidence42
The $15.5bn valuation is 38.75 times the $400m of annual recurring revenue Harvey disclosed, and the in-house model meant to defend that price is post-trained on a Beijing lab's open weights rather than on backer OpenAI's.
Reality
- Evidence38
- Adoption62
- Hype gap+34
- Incentives72
- Confidence44
Ramp's Q4 2025 spend data sizes the model hosting and serving layer at $260m across roughly 1,900 buyers, which is 60 cents for every dollar those same companies hand straight to OpenAI and its closed-source peers.
Publishers:ramp.com
Reality
- Evidence55
- Adoption42
- Hype gap−10
- Incentives55
- Confidence52
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
Reality
- Evidence32
- Adoption42
- Hype gap+34
- Incentives88
- Confidence44