Build1 distinct publisher2 min readPublished
A proposed scaling law for distillation says the winner flips on two things: whether you already own a teacher, and how many students you intend to serve.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Teacher compute behaves like a fixed cost and student compute like a variable one, which is why the paper's two recipes point in opposite directions [2]. Train a teacher to serve a single student and the whole teacher bill lands on that one model; inherit a teacher someone else paid for, or spread it across a roster, and the charge per student falls toward nothing [15]. The reported result that distillation wins with many students or an existing teacher [3], and loses when there is one student and a teacher still to train [4], is the same statement read from both ends of that division [15].
The motivation is the serving bill, not the training one. Compute-optimal models get larger as budgets grow, which is what makes them awkward to deploy [5], and the paper cites work putting an LM's inference cost typically well above its pretraining cost [6]. The alternative already in use is overtraining: far more data than compute-optimal, which yields small capable models [7] but pushes them into diminishing returns and many trillions of tokens to get there [8]. Distillation is offered as a way to reach the same small-model capability for less training spend [9].
The ceiling is the part worth pinning down. The advantage holds up to a compute level that scales predictably with student size [3], which makes the crossover a property of the student you picked rather than of the budget you approved [16]. The team most exposed to it is the one following the overtraining playbook: pushing a fixed small model far past compute-optimal token counts [8][16].
There is a reason to be careful about extrapolating the curve. The authors state plainly that there is no consensus on why distillation works, listing dark knowledge transfer, regularization and noise reduction among the competing explanations [11]. A fitted law over an unexplained effect is dependable inside the range it was fitted on and a guess outside it.
The abstract also presents allocation as risk mitigation [14] without printing the constants. "Scales predictably with student size" is a shape, not a number, and neither the abstract nor the introduction supplied here gives the value [17]. Anyone planning to run this arithmetic before committing a cluster has to lift the fitted coefficients out of the paper body and check that their own student sizes and token budgets sit inside the studied range. That is a modest ask, and it is what the contribution actually is: the choice between distilling and training becomes a calculation with inputs a finance function already tracks [1].
Ranked by verification strength, evidence, and original report placement.
The paper proposes a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher.
The paper provides compute-optimal distillation recipes for two scenarios: when a teacher already exists, and when a teacher needs training.
In settings involving many students or an existing teacher, distillation outperforms supervised learning up to a compute level that scales predictably with student size.
If only one student is to be distilled and a teacher also requires training, supervised learning is generally preferable.
The size of compute optimal models grows with compute, which makes them challenging to use because of growth in inference costs.
The paper states that the inference cost of a language model is typically significantly larger than its pretraining cost, citing Chien et al. 2023 and Wu et al. 2024a.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Large controlled sweep, single-lab preprint, key numbers absent here
The claims trace to a described large-scale controlled study (transformer students and teachers from 143M to 12.6B parameters, data from a few billion to 512B tokens) and the paper states its conditions and limits rather than only its wins. What holds the score down is that both supplied sources are the same preprint, there is no independent replication, and the supplied abstract and introduction contain no crossover compute value or fitted coefficients, so the central quantitative claim cannot be checked from this material.
Method already in shipped model families; the law itself unadopted
Distillation pretraining has demonstrable production footprint: the paper reports its use in the Gemma/Gemini, Minitron and AFM families. Adoption of the contribution at issue, the scaling law and its two recipes, has no evidence at all in the supplied sources beyond the preprint's own publication, and the production evidence is secondhand citation rather than direct disclosure, with one cited counter-result.
Slightly overstated in the abstract, corrected in the introduction
The abstract's 'mitigate the risks associated with large-scale distillation' and its clean two-recipe verdict read stronger than what the supplied text substantiates, since the crossover level and coefficients are not given here and the advantage is explicitly bounded. The gap is small because the same paper volunteers the limits: no consensus on the mechanism, a cited contrary result, and the statement that distillation cannot beat supervised cross-entropy given enough data or compute.
No affiliation, funding or interest disclosure in supplied text
The supplied abstract page and article body contain no author names, institutional affiliations, funding statements or competing-interest declarations, and no commercial product is being sold in the sources. Any incentive reading would require inferring sponsorship from cited model families, which the material does not support.
Direction credible, magnitude and independence unverified
Confidence is moderate: the decision logic is internally consistent, stated twice by the primary source, and consistent with the amortization reading, and the technique has real production usage. It is capped by single-publisher sourcing, absent quantitative thresholds, no replication, an unresolved contradiction in prior work, and no incentive disclosure.
product
Model choice is becoming a line item, and the differentiator moved up the stack1 distinct publisher
leadership
Google put the model in the car: a Pixel on the CAN bus is the deployment shape nobody budgeted1 distinct publisher
build
The judge went synthetic first, which tells you which part of your pipeline is next1 distinct publisher
security
Google's reference agent approved a $10,000 refund on a $149 order, on purpose1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · August 26, 2026