Product1 distinct publisher2 min readUpdated
Llama 3.1 was pretrained on 15 trillion tokens. A preteen manages on about 100 million words. Nobody knows how the child does it, and easily available web data may run out in the 2030s.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The multiple is worth doing by hand. Technology Review puts the difference at roughly a hundred thousand times more words for a model than for a person mastering a first language [11]; the published figures tighten that. Llama 3.1's 15 trillion pretraining tokens [2] against a well-read preteen's 100 million words [4] is about 150,000 to one [14]. If Wilcox is right that frontier runs use ten times more data [3], the ratio is nearer 1.5 million to one [15]. Tokens are word-like chunks rather than words [13], so read these as orders of magnitude, not measurements. They do not get smaller at any level of precision.
What makes the ratio a problem rather than a curiosity is where the low end sits. A toddler is producing grammatically correct sentences after something like 10 million words, 30 million at the outside [5]. That is roughly six orders of magnitude below the last publicly stated frontier-scale pretraining volume [16]. The comparison is not that machines are behind children; it is that the two systems are not on the same curve, and only one of them has a known method.
For the past decade, language models have mostly improved by getting bigger [17]. The alternative, learning more from less, has no owner and no schedule, because exactly how babies do it remains unknown [8]. It is not even settled whether a child arrives with a language instinct or could in principle learn a language from experience alone [9]. That second uncertainty is the one operators should weigh. If part of the child's efficiency is biological rather than architectural, then "a more data-efficient learner" is not an engineering task waiting for headcount. It is a bet on a result that cognitive science and machine learning have both failed to produce so far.
The field's own name for this, the data efficiency gap [1], flatters the situation. A gap implies a measured distance and a crossing. What the researchers describe is closer to an absent component: the only learner that reaches fluency on a household's worth of language cannot be inspected well enough to copy, while the learner that can be inspected needs a city's worth of text and then some [8].
Anyone treating data volume as the dependable input to capability is relying on a supply with an estimated end date [7] and no demonstrated successor. That is not a reason to discount current models, which work. It is a reason to stop describing the next order of magnitude as progress toward efficiency, when it is the opposite of it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Cognitive scientists call the divide between how much language children and machines need in order to learn the data efficiency gap.
Meta's open-weight LLM Llama 3.1, released two years before the article, chewed through 15 trillion tokens in pretraining.
A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words; adding literacy can push the count to maybe 300 million words by age 20.
Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words, or 30 million on the high end.
Michael C. Frank, a cognitive scientist at Stanford University, says: "If you train GPT-2 on 30 million words, you get a nonsense generator; you don't get a kid."
Exactly how babies learn language is a mystery; researchers know a lot about what kids learn and how they use language at different stages, but much remains unknown.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, named experts, order-of-magnitude estimates
All evidence comes from a single MIT Technology Review article. Its strength is attribution: figures and interpretations are sourced to named academics (Frank at Stanford, Wilcox at Georgetown, Futrell at UC Irvine) and one published model figure (Llama 3.1's 15 trillion tokens). Its weakness is that the child-exposure counts are explicitly hedged approximations, no primary papers or datasets are cited in the supplied text, and the two most consequential quantitative claims -- 10x frontier data and 2030s data exhaustion -- rest on a single unverified estimate each.
Scale is the deployed path; data efficiency is lab-stage
Adoption evidence points the opposite way from the story's aspiration. What is demonstrably in production is large-scale pretraining: Llama 3.1 shipped on 15 trillion tokens, and the article says a decade of improvement came mostly from getting bigger. The data-efficient alternative appears only as a research goal with named prospective uses (video training, minority-language chatbots) and one cited negative result -- GPT-2 on 30 million words produces nonsense. No supplied source shows a data-efficient model in deployment or benchmarked against frontier systems.
Slightly overstated framing on hedged numbers
The article itself is restrained: it hedges the child-exposure counts, attributes the frontier multiplier, and qualifies the data-exhaustion date with "perhaps as early as." The overstatement is in the crispness that the framing and derived ratios lend to soft inputs -- a clean 150,000-to-1 and 1.5-million-to-1 read as measurements when they are arithmetic on ranged estimates that also compare tokens to words as if they were the same unit. The headline claim that scaling is a workaround for an undiscovered mechanism is supported by the article's own concession of ignorance, so the gap is small and positive rather than large.
Academic sources, no vendor promotion
The observable incentive structure is comparatively clean: every named voice is an academic cognitive scientist or linguist, no company is selling a product in the story, and the only corporate figure cited (Meta's Llama 3.1 token count) is used as third-party evidence rather than a vendor claim. The residual incentive is disciplinary -- researchers quoted on the importance of reverse-engineering child learning benefit professionally if that agenda gains standing, and the supplied source discloses no funding or competing interests. There is no evidence of commercial or promotional incentive in the supplied material.
Moderate: well-attributed but single-source and hedged
Confidence is moderate. The descriptive core -- a multi-order-of-magnitude data gap, and no known mechanism explaining child learning -- is consistently and credibly attributed, and one anchor figure is a published model detail. But the cluster has a single publisher, the child-side numbers are ranged estimates, and the forward-looking elements (frontier 10x, 2030s exhaustion) are unverified. That supports the story's framing while leaving its most decision-relevant numbers weakly grounded.
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
invest
Anthropic's discounts leave a 7.5x-to-37.5x gap, and the routing code decides the rest1 distinct publisher
security
The nationalization argument is really a vendor-continuity memo1 distinct publisher
leadership
Builders put doom at 10 to 50 per cent and expect binding rules only after the disaster1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026