Product1 publisher3 min readPublished
Sutton calls synthetic data 'a big mistake': every simulator is a lossy copy of a bigger world
The reinforcement learning co-founder told Sequoia's podcast that no simulation can cover other minds or the physical world. The labs shipping synthetic corpora are aiming at the domains he exempts.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Richard Sutton shared the 2024 Turing Award with Andrew Barto for founding reinforcement learning.
- Sutton said of the turn to synthetic data, "That's just a big mistake," on Sequoia Capital's Training Data podcast, published Tuesday and hosted by Sonya Huang and Pat Grady.
- Epoch AI projects that the stock of public human text, somewhere around 300 trillion tokens, will be fully used between 2026 and 2032.
- AI companies have been buying and cutting up second-hand books for pre-2022 text, on the reasoning that anything published since is contaminated with machine-generated slop.
- Chinese labs are hitting the same data shortage in their own language, and rather sooner.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Richard Sutton, who shared the 2024 Turing Award with Andrew Barto for founding reinforcement learning, called the industry's turn to synthetic data "just a big mistake" on Sequoia Capital's Training Data podcast, published Tuesday and hosted by Sonya Huang and Pat Grady [1][2]. The remark is aimed at the workaround most current roadmaps assume, because Epoch AI projects that the stock of public human text, somewhere around 300 trillion tokens, will be fully used between 2026 and 2032 [3].
The squeeze is physical as well as statistical. AI companies have been buying and cutting up second-hand books for pre-2022 text, on the reasoning that anything published since is contaminated with machine-generated output [4], and Chinese labs are hitting the same wall in their own language rather sooner [5].
Sutton's objection has two specific edges rather than being a blanket verdict [6]. "There's no way we can have synthetic data for other people's minds," he said, and on simulating the physical world: "The world is infinitely complex, and any simulation of it is like, microscopic" [7][8]. Underneath both sits what he calls the big world hypothesis: reality is always larger than an agent's model of it, which makes any simulator a lossy compression by construction [9]. The second edge is procedural. Somebody has to decide which synthetic data to generate, and that smuggles human judgement back into exactly the process his 2019 essay The Bitter Lesson warned against [10][11]. What he wants instead is experiential data, gathered by an agent acting in its environment rather than assembled in advance by anyone [12]. The thesis is not new, having been set out with David Silver in Welcome to the Era of Experience in April 2025, but he has not previously aimed it this squarely at synthetic data [13].
The published record cuts both ways. A 2024 Nature paper by Ilia Shumailov and colleagues found that models trained recursively on their own output degrade, an effect now called model collapse [14]. Work by Matthias Gerstgrasser, Rylan Schaeffer and co-authors found that collapse depends on synthetic data replacing human data, and does not show up where the two accumulate alongside each other [15]. Replacement versus accumulation is the distinction an operator can actually act on.
Meanwhile the labs have shipped. Microsoft's Phi-4 trained on roughly 400 billion synthetic tokens across 50 dataset types [16], and Nvidia has released a synthetic pre-training corpus of about 10 trillion tokens for its Nemotron models [17], roughly 3 percent of Epoch's estimate for all public human text [18]. Those are also the domains where Sutton's objection bites least, since maths and code have checkable answers in a way that other people's minds do not [19]. His limits are about modelling humans and physical systems, not about generating training examples as such [20].
Andrej Karpathy has made the best-known counter-case, arguing that language models are less like animals grown through experience than ghosts distilled from human writing, and that tuning them is a legitimate path rather than a wrong turn [21]. Sutton told the podcast that language accounts for perhaps a quarter of intelligence [22].
Sutton is calling this the next big lesson, and Sequoia is billing it as a second bitter one, tidy framing from a firm that also backs Silver's new lab [23]. Sutton left John Carmack's Keen Technologies in July to start his own [24], co-founded with his former student Khurram Javed and aiming at a trillion-parameter mind that never stops learning and runs on 20 watts, within five to ten years [25]. For anyone budgeting data this year, the testable question is narrower than the headline: whether the synthetic tokens accumulate next to human ones or displace them, and whether the target domain grades its own answers.