Build1 distinct publisher3 min readUpdated
Hugging Face now hosts 2.96 million public model repositories. The ones anyone actually pulls number in the tens of thousands, and the monthly parameter ceiling has been set in China all year.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Hugging Face's Summer 2026 hub review reports public model repositories growing from 2.43 million to 2.96 million over the period, with datasets going from 711,000 to 1 million and Spaces from 1.00 million to 1.44 million [1]. In the same post, 1.5% of repositories account for 99.2% of all downloads [2], which means the count that gets quoted in vendor decks and the count that matters to anyone picking a model differ by roughly two orders of magnitude.
Run the arithmetic on Hugging Face's own percentages and the shortlist gets concrete. Applied to the 2.96 million model repos, 1.5% is about 44,000 repositories that absorb essentially all traffic [1]. The 85.6% of models with fewer than 200 lifetime downloads [3] works out to roughly 2.53 million repositories that are, functionally, archived files [2].
At the top of the size range the split is geographic. Hugging Face reports that in almost every month of 2026 the largest and most performant open model from a Chinese lab was bigger than anything an American lab released of its own, with China's monthly ceiling running between 754 billion and 2.78 trillion parameters [4]. The American ceiling stayed under 130 billion in five of seven months, the exceptions being NVIDIA's Nemotron 3 Ultra at 561 billion in May and June and Inkling from Thinking Machines Lab [5]. At the extremes that is about a 21x gap in parameter count [3].
The portfolio shapes diverge too. Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70 billion parameters, so a developer's first encounter is a model too large to run locally [6], while Tencent and Alibaba Qwen cover the range from under 1 billion upward [7]. Xiaomi and Meituan both cleared a trillion parameters this year [8]. The frontier-only strategy rests on a dependency worth naming: Hugging Face argues a lab no longer has to ship a small model to be reachable, because the community quantization layer makes a large one runnable within days [9].
American participation has moved desks rather than disappeared. The two organisations publishing the most new open models this year are AMD and NVIDIA, each with more than 200 new model repositories, with LiquidAI third at around 100 [10]. Google and Meta now rank well below NVIDIA in new releases [11], and Meta has moved toward closed flagship models [12]. Above 100 billion parameters, Hugging Face says most U.S. releases this year are not new models but are built on top of Chinese ones [13]; the named originals are Inkling at 952B, Nemotron 3 Ultra at 561B, Nemotron 3 Super at 124B, and Arcee AI's Trinity-Large at 399B [14]. AMD contributed many conversions and no original model at that scale [15]. Chinese open models, meanwhile, are increasingly optimised for domestic chips [16].
One more number for anyone using engagement as a proxy for adoption: of the top 25 repos by downloads this year and the top 25 by likes, exactly one appears on both lists [17]. No model published in 2026 reaches the download top 25, and thirteen of the twenty-five date from 2022 [18]. all-MiniLM-L6-v2 was pulled 1.55 billion times in seven months against 5,156 likes [19], about 300,000 downloads per like [4], against roughly 60 per like for Kimi-K3 [20] - a ratio gap near 5,000x [5].
Watch whether any 2026 release breaks into the download top 25, whether AMD publishes an original model above 100B rather than conversions, and whether the quantization layer keeps absorbing trillion-parameter releases fast enough to justify frontier-only portfolios.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In almost every month of 2026 the largest and most performant open model from a Chinese lab was larger than anything an American lab released of its own; China's monthly ceiling ran between 754 billion and 2.78 trillion parameters.
America's own monthly parameter ceiling stayed under 130 billion in five of seven months, the exceptions being NVIDIA's Nemotron 3 Ultra at 561 billion parameters in May and June, and Inkling from Thinking Machines Lab.
Google and Meta now rank well below NVIDIA in new model releases, despite having defined open model publishing in previous years.
all-MiniLM-L6-v2 was pulled 1.55 billion times in seven months against 5,156 likes.
Kimi-K3 was pulled about 60 times per like it received.
Public model repositories on the Hugging Face hub grew from 2.43 million to 2.96 million over the period covered; datasets grew from 711,000 to 1 million; Spaces from 1.00 million to 1.44 million.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified but single-source and unreproducible
Every figure comes from one publisher that is also the platform being measured, using telemetry no one else can query. The quantified core is unusually specific and internally consistent — repository counts, concentration ratios, windowed download leaders, per-account heavy-band shares, named models with parameter counts — and the post volunteers a methodological control (downloads counted inside the window, not lifetime) plus a self-correction on likes-versus-downloads. Against that: no data export or reproducible query, no exact window dates, 'most performant' is asserted without benchmark scores, several load-bearing claims (quantization within days, domestic-chip optimisation, Meta's closed pivot) carry no supporting data, and the supplied body is truncated mid-argument.
Heavily documented usage, extremely concentrated
Adoption is directly measured rather than inferred: disclosed inventory growth, per-account download shares, publisher totals (Moonshot 37M versus Qwen 2,045M), and a single embedding model at 1.55 billion pulls in seven months. Release-side adoption is also concrete, with AMD and NVIDIA each above 200 new repositories and named frontier models at both ceilings. The score is held below the top band because the same data shows adoption is confined to roughly 1.5% of repositories, and because the frontier models that set the ceiling attract likes far more than pulls — high publication volume is not the same as deployed usage.
Mostly deflationary, with parameter count doing overstated work
The story is largely hype-reducing: it argues headline repository counts are meaningless, that 1.5% of repos take 99.2% of downloads, that likes and downloads measure different things, and it flags the publisher's own earlier coverage as having made that mistake. That pushes the gap toward zero. Residual positive gap comes from the framing that 'the ceiling is Chinese', which treats parameter count as a stand-in for capability without a single benchmark score, and from unquantified supporting assertions about quantization turnaround, domestic-chip optimisation and Meta's closed pivot that the narrative leans on.
Platform operator reporting on its own platform
The sole source is Hugging Face writing about the Hugging Face hub, using metrics only it can produce, and its commercial position benefits from open-weights publishing volume and hub centrality. The post also serves as an argument that the community layer around the hub (quantization, conversions) is what makes frontier weights usable. Mitigating factors keep this from the top of the range: the analysis actively deflates the hub's own headline growth numbers, discloses an error in its prior coverage, and reports unflattering distribution facts. No pricing, licensing or sponsorship disclosure appears in the supplied text.
Single-source, self-measured, partially truncated
One publisher, one document, no corroboration or contradiction available anywhere in the cluster, and that publisher is the measured party. Confidence is lifted above the floor by the specificity and internal consistency of the numbers, the stated age control on downloads, and the publisher's privileged access to the underlying data; it is capped by the absence of any second vantage point, missing methodology detail, a truncated body, and several claims in the ledger that the supplied text only asserts.
product
Cheap bug-hunting arrives: GLM 5.3 puts near-frontier vulnerability discovery on your own hardware1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
invest
The chips never move: Washington's fix for the Southeast Asia compute loophole1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 13, 2026