Invest1 distinct publisher3 min readUpdated
A new paper argues Chinchilla treats model size and data as independent when they are not, and that fixing it lets teams fit the same scaling curve on a sparse grid.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Meta's FAIR lab has published a paper arguing that DeepMind's Chinchilla scaling law mis-specifies the relationship between model size and training data, and that a corrected version reaches near-identical accuracy while consuming roughly 10x less compute to fit [1][3][8]. The money at stake is not in the headline training run but in the grid of small profiling runs teams burn before it, which on this account has been sized about an order of magnitude larger than necessary [8][13].
The mechanism is unglamorous. Chinchilla, established by DeepMind in March 2022, treats parameters (N) and training tokens (D) as independent variables that each contribute separately to final loss [2]. The Skaling authors computed the mixed partial derivative of the loss surface with respect to N and D and report that it is non-zero [4]: the payoff from adding parameters depends on how much data you have, and the reverse. Their proposed form wraps the Chinchilla expression in an exponent, L(N, D) = (A / N^a + B / D^b)^k + E, where k greater than 1 encodes the coupling [3]. Set k to 1 and you get Chinchilla back exactly, which means the old law is a special case rather than a competitor [3][14].
The reported margins, according to Cryptobriefing's account of the paper, are wide but not universal: Skaling beat Chinchilla at 76% of tested configurations, with a median 2.2x improvement in prediction accuracy, and improvements of 4x or better at a third of tested points [5][6]. Mean absolute percentage error fell by 1.5 to 3x across scenarios [7]. That also leaves 24% of configurations where it did not win [15]. Note what is being measured: accuracy in predicting loss, not better models.
The line item is the profiling grid. Skaling is reported to hold its accuracy on a sparse "L-shaped" grid, where you vary one dimension with the other fixed and then reverse, instead of the full grid Chinchilla requires, at roughly 10x less compute [8]. If that survives contact with other labs, the implication for anyone budgeting model development is that about 90% of the compute historically allocated to scaling experiments was buying precision that a two-arm sweep already delivers [8][13].
It matters most for the way models are actually shipped. Chinchilla's 2022 contribution was to show that many models were badly undertrained [10], but production models are now routinely overtrained well past its optimum because inference economics reward small models fed enormous amounts of data [11]. The coupling exponent is presented as capturing how the optimal balance moves in exactly that regime [12], and at frontier compute levels the paper reports optimal token-to-parameter ratios that can differ from Chinchilla's by orders of magnitude [9].
What to watch: independent refits by teams with their own loss surfaces, since the reported result is a ratio and the account gives no absolute compute figures or fitted values of k [15]. Watch whether the 10x saving holds when the L-shaped grid is used to extrapolate to a run much larger than any point in it, which is the only use that matters. And watch the disclosed token-to-parameter ratios of models released over the next two quarters; if the coupling correction is real, those ratios should start moving away from Chinchilla-derived defaults [9].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Skaling outperformed Chinchilla at 76% of tested configurations, with a median improvement factor of 2.2x in prediction accuracy.
At one third of the tested points, Skaling's improvement over Chinchilla reached 4x or higher.
The new law reduced mean absolute percentage error by 1.5 to 3x compared with Chinchilla across various test scenarios.
Skaling achieves nearly identical accuracy using a sparse "L-shaped" profiling grid, where one dimension is varied while the other is held fixed and then the procedure is repeated in the other direction, consuming approximately 10x less compute than the full-grid method Chinchilla needs.
The paper "Skaling: Chinchilla's Exponents Meet Kaplan's Coupling" was published on August 7, 2026 by Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz and Kartik Ahuja of Meta's FAIR lab.
DeepMind's Chinchilla scaling law, established in March 2022, treats model parameters (N) and training tokens (D) as independent variables that each contribute separately to a model's final loss.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single secondary account of an unlinked paper
Every claim traces to one article from one publisher, which restates figures attributed to the paper without a link, quotation, replication or peer-review status. The reported quantities are specific and internally consistent (76% win rate, 2.2x median, 1.5-3x MAPE reduction, ~10x fitting-compute reduction) and the functional form is stated precisely enough to test, which lifts this above rumour. But the frontier-ratio and overtraining-guidance statements carry no supporting numbers at all, and nothing in the cluster independently verifies any figure.
No usage evidence beyond publication
The only observable event is the publication of the paper itself. The supplied material reports no lab, team or toolchain fitting the Skaling law or using the L-shaped profiling grid, no released code or artefacts, and no training run planned with it. A paper release is not adoption, and inferring uptake would require facts the sources do not contain.
Cost claim generalised beyond what it covers
The headline and subhead frame the result as cutting compute costs by 10x, and the cluster title says the standard way to plan a training run costs 10x more than it needs to. The underlying finding is narrower: roughly 10x less compute for the scaling-law profiling sweep at near-identical fit accuracy, which is a planning-stage overhead rather than the training run. The gap widens because the 24% of configurations with no improvement goes unmentioned, the frontier-ratio and overtraining-guidance claims are asserted without figures, and no adoption exists yet. It is not a fabricated result - the mechanism and numbers are concrete and Chinchilla is genuinely nested at k=1 - so the overstatement is scope inflation, not invention.
Interested-party research plus off-beat amplifier
The paper's authors are employees of Meta's FAIR lab and the result reframes a rival lab's (DeepMind's) widely used scaling law as a special case of Meta's, which is a reputational and standards-setting benefit to the publishing institution; the supplied material gives no independent evaluation. The amplifying outlet is a cryptocurrency publication covering a machine-learning methods paper, a placement whose interest is traffic on an AI cost headline rather than methodological scrutiny. No funding, pricing, licensing or commercial-conflict facts are supplied, so this reads institutional incentive rather than a documented financial one.
Low - one outlet, no verification, no uptake
Confidence is limited by the cluster's structure rather than by contradiction: a single publisher, no primary-paper access, no replication, no adoption signal, and two of the more consequential statements unsupported by figures. Nothing in the material contradicts the claims either, so the assessment is thin rather than disputed. The precise functional form and the k=1 nesting are the parts most likely to hold on later verification.
product
Callosum raises $100m for mixed-silicon scheduling, and the 2x accuracy claim is still its own2 distinct publishers
product
Vivodyne says the AI drug bottleneck is human tissue data, not model capability1 distinct publisher
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
build
Half of Claude's watermark ships with a reference tool. Your PDF pipeline eats it.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
cryptobriefing.com
1 article · August 15, 2026