Product1 distinct publisher3 min readUpdated
The reinforcement learning co-founder told Sequoia's podcast that no simulation can cover other minds or the physical world. The labs shipping synthetic corpora are aiming at the domains he exempts.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Richard Sutton, who shared the 2024 Turing Award with Andrew Barto for founding reinforcement learning, called the industry's turn to synthetic data "just a big mistake" on Sequoia Capital's Training Data podcast, published Tuesday and hosted by Sonya Huang and Pat Grady [1][2]. The remark is aimed at the workaround most current roadmaps assume, because Epoch AI projects that the stock of public human text, somewhere around 300 trillion tokens, will be fully used between 2026 and 2032 [3].
The squeeze is physical as well as statistical. AI companies have been buying and cutting up second-hand books for pre-2022 text, on the reasoning that anything published since is contaminated with machine-generated output [4], and Chinese labs are hitting the same wall in their own language rather sooner [5].
Sutton's objection has two specific edges rather than being a blanket verdict [6]. "There's no way we can have synthetic data for other people's minds," he said, and on simulating the physical world: "The world is infinitely complex, and any simulation of it is like, microscopic" [7][8]. Underneath both sits what he calls the big world hypothesis: reality is always larger than an agent's model of it, which makes any simulator a lossy compression by construction [9]. The second edge is procedural. Somebody has to decide which synthetic data to generate, and that smuggles human judgement back into exactly the process his 2019 essay The Bitter Lesson warned against [10][11]. What he wants instead is experiential data, gathered by an agent acting in its environment rather than assembled in advance by anyone [12]. The thesis is not new, having been set out with David Silver in Welcome to the Era of Experience in April 2025, but he has not previously aimed it this squarely at synthetic data [13].
The published record cuts both ways. A 2024 Nature paper by Ilia Shumailov and colleagues found that models trained recursively on their own output degrade, an effect now called model collapse [14]. Work by Matthias Gerstgrasser, Rylan Schaeffer and co-authors found that collapse depends on synthetic data replacing human data, and does not show up where the two accumulate alongside each other [15]. Replacement versus accumulation is the distinction an operator can actually act on.
Meanwhile the labs have shipped. Microsoft's Phi-4 trained on roughly 400 billion synthetic tokens across 50 dataset types [16], and Nvidia has released a synthetic pre-training corpus of about 10 trillion tokens for its Nemotron models [17], roughly 3 percent of Epoch's estimate for all public human text [18]. Those are also the domains where Sutton's objection bites least, since maths and code have checkable answers in a way that other people's minds do not [19]. His limits are about modelling humans and physical systems, not about generating training examples as such [20].
Andrej Karpathy has made the best-known counter-case, arguing that language models are less like animals grown through experience than ghosts distilled from human writing, and that tuning them is a legitimate path rather than a wrong turn [21]. Sutton told the podcast that language accounts for perhaps a quarter of intelligence [22].
Sutton is calling this the next big lesson, and Sequoia is billing it as a second bitter one, tidy framing from a firm that also backs Silver's new lab [23]. Sutton left John Carmack's Keen Technologies in July to start his own [24], co-founded with his former student Khurram Javed and aiming at a trillion-parameter mind that never stops learning and runs on 20 watts, within five to ten years [25]. For anyone budgeting data this year, the testable question is narrower than the headline: whether the synthetic tokens accumulate next to human ones or displace them, and whether the target domain grades its own answers.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Richard Sutton shared the 2024 Turing Award with Andrew Barto for founding reinforcement learning.
Sutton said of the turn to synthetic data, "That's just a big mistake," on Sequoia Capital's Training Data podcast, published Tuesday and hosted by Sonya Huang and Pat Grady.
Epoch AI projects that the stock of public human text, somewhere around 300 trillion tokens, will be fully used between 2026 and 2032.
Sutton's objection has two specific edges rather than being a blanket verdict.
Sutton said: "There's no way we can have synthetic data for other people's minds."
On simulating the physical world, Sutton said: "The world is infinitely complex, and any simulation of it is like, microscopic."
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Well-attributed core, single outlet, two unsourced assertions
The central claims are traceable: direct podcast quotes, a named 2024 Nature paper and its named published rebuttal, an Epoch AI projection, and specific token counts for Phi-4 and the Nemotron corpus. But the cluster contains exactly one source article, no primary links, no comment from the labs or researchers named, and at least two load-bearing colour claims (second-hand book scanning, Chinese-language shortage arriving sooner) that are asserted without attribution.
Synthetic corpora already in production; experiential alternative unshipped
Adoption of the practice Sutton criticises is concrete and disclosed: roughly 400 billion synthetic tokens across 50 dataset types in Phi-4, and a released ~10 trillion token synthetic pre-training corpus for Nemotron, plus reported physical scavenging of pre-2022 text. Adoption of Sutton's alternative is effectively nil in the supplied material - his lab is newly formed and its targets are five to ten years out - so the measured figure describes the synthetic-data build-out, not the thesis being argued for.
Blunt verdict outruns the narrow evidence, though the article self-corrects
The framing - a Turing laureate declaring the industry's data fix 'a big mistake' - is broader than what the cited evidence sustains. The peer-reviewed collapse result applies to recursive self-training and is narrowed by published work showing no collapse when human and synthetic data accumulate together, and the corpora actually shipped sit in maths and code, the domains Sutton's own objection exempts. The gap is moderate rather than large because the article itself carries the rebuttal, the exemption, Karpathy's counter-case and the disclosure of Sequoia's stake in an adjacent lab.
Venture-hosted venue, disclosed adjacent stake, and a founder promoting his own thesis
The claim originates on a venture firm's own podcast; that firm is billing the thesis as a second bitter lesson and, as the article discloses, backs David Silver's new lab, Silver being Sutton's co-author on the experiential-learning essay. Sutton himself left Keen Technologies in July to found a lab whose entire premise is the alternative he is advocating. The commercially interested parties on the other side - Microsoft and Nvidia, whose synthetic corpora are cited - are not quoted. Incentives are visible and disclosed rather than hidden, but they are dense enough to shape the framing.
Single publisher, verifiable quotes, contested underlying science
Confidence is capped by the cluster containing one article from one outlet with no primary links. What raises it above the floor is that the quotations are attributable to a dated public podcast, the two research findings and the two corpus figures are specific and independently checkable in principle, and the article discloses its own framing conflict. What holds it down is the absence of corroboration, the unattributed sourcing claims, and the fact that the central question - whether synthetic data is a mistake - is a matter of contested science rather than settled record.
leadership
Microsoft has 2.2m AI chips installed. Its own capacity claims imply up to 6.4m1 distinct publisher
product
Nebius funds $4.5bn of AI capacity on terms that pay lenders mostly in stock2 distinct publishers
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
Nvidia is brokering the Nordic build-out, not just supplying it2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026