BuildNot yet confirmed elsewhere1 publisher3 min readPublished
The judge went synthetic first, which tells you which part of your pipeline is next
Latent Space dates five flips from human-made to model-made since 2022. The order they happened in predicts more than the annual cadence the essay claims for it.
The Engineer · Build desk

What happened
- Latent Space dates one component of the intelligence pipeline flipping from human-made to model-made each year since 2022, each with a patient-zero release at a frontier lab.
- The corpus followed in 2023, including Apple's WRAP, which rephrased the web with an LLM and made pretraining roughly 3x more efficient.
- Stage five is the researcher: Karpathy's March 2026 autoresearch loop edits a training setup, runs a five-minute experiment, and keeps the change only if validation loss drops.
Why it matters
- decision It changes where an operator looks for the next substitution: the touchpoint humans are asked to service most often per iteration, not the one with the biggest invoice.
- constraint The 100x discount only converts into quality where you can run and select among many attempts, which rules out any step with one shot and no cheap verifier.
- exposure Once distillation is the default assumption for small model releases, shipping a strong reasoner also supplies the training signal for whoever wants to copy it.
- contradiction The annual framing invites you to expect a flip on schedule, while the dates behind it cluster and skip, so this is a diffusion pattern rather than a clock to plan against.
The ordering is the part worth arguing about. If cost alone decided, the corpus would have gone first: text is the largest line item and the easiest thing to imitate. Instead the judge went first, in 2022, when InstructGPT collected human preferences once and let the policy optimise against a trained reward model rather than against people [3]. A reward signal is consulted once per sample. A pretraining corpus is assembled once per run. The stage that flipped earliest was the one where humans sat in the loop the most times per unit of progress [19], and once LLM-as-judge became the default evaluation method, approval and scoring both ran model-on-model [5].
The price of that substitution is stated plainly in the essay: roughly 10% worse, 100x cheaper, 10,000x faster [2]. At a hundredth of the cost, the same budget buys about a hundred attempts instead of one [17]. That arithmetic only pays where you can select or verify among the attempts. Where you get one shot and no cheap check on the output, a 10% quality penalty is just a 10% quality penalty.
Which is why the last two flips in the sequence are the interesting ones. Curriculum design, described here as historically the most artisanal and taste-driven part of ML [12], flipped in 2024 when self-rewarding models and SPIN showed a model generating its own tasks, grading its own outputs and improving past the ceiling of its human preference data [11]. Stage five, the researcher, arrives with Karpathy's autoresearch in March 2026: a coding agent edits a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight [15]. In an eight to twelve hour window that is roughly 96 to 144 sequential experiments [18].
Note what makes that loop possible. The acceptance test is validation loss, a metric that costs nothing to consult. Stage five is parasitic on stage one: the researcher can only be automated because the judge already was. That is the useful part of the framework for anyone auditing their own stack. The candidate for the next flip is not the most expensive human input, it is the human input you consult most often per iteration and can replace with something a machine can check. If your acceptance test is taste with no metric behind it, the flip stalls, which is exactly why curriculum needed a self-reward signal to arrive first.
Two cautions. This is one publisher's mental model, offered as something you arrive at if you squint at the same reading list [1] [2], and its own dates do not march to an annual beat: two flips land in 2023, none in 2025 [16]. And the essay groups end-to-end RL environments with synthetic data and synthetic rubrics as the same family of increasingly ambitious human simulation [2], while the five numbered stages we have run reward [3], data [6], teacher [9], curriculum [11] and researcher [13]. Environments get named as the pattern and then not dated. On the framework's own logic that is a stage waiting for its patient zero, not a stage that skipped its turn.
What to watch
- Whether a ratchet-style autoresearch loop shows up inside a frontier lab's published training pipeline rather than a personal repo.
- Whether end-to-end RL environments get a dated patient-zero release, which is the stage the framework names but does not place.
- Whether anyone publishes the cost and quality measurements behind the 10% worse, 100x cheaper claim instead of asserting it.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence42
- Adoption58
- Hype gap+32
- Incentives55
- Confidence48
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Latent Space argues that every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made, and that each flip has a patient zero: a paper or product where the synthetic version first became load-bearing at a frontier lab, after which it diffuses.
- [2]
The essay frames synthetic data, synthetic rubrics, the AI researcher and end-to-end RL environments as one family of increasingly ambitious human simulation that is roughly 10% worse, 100x cheaper and 10,000x faster.
- [3]
Stage 1, the reward signal, went synthetic in 2022: InstructGPT established collecting human preferences once, training a reward model, and letting the policy optimise against the model rather than against humans.
- [4]
Constitutional AI had the model critique itself against a set of principles (RLAIF), and Lee et al. later showed AI feedback matching human feedback at a fraction of the cost.
- [5]
LLM-as-judge became the default evaluation methodology via MT-Bench and AlpacaEval, at which point reward, critique and evaluation all ran on models judging models.
- [6]
Stage 2, the training data, went synthetic in 2023: Microsoft's Phi series trained a small model on LLM-synthesised textbook-quality data that punched above its parameter count, and phi-1.5 confirmed it was not a fluke.
- [7]
Apple's WRAP generalised the move by rephrasing the entire web with an LLM, making pretraining roughly 3x more efficient.
- [8]
NVIDIA's Nemotron-4 340B shipped a permissively licensed synthetic data generation pipeline as a headline feature, and by 2025 reasoning-trace corpora generated by strong reasoners had become a standard pretraining and mid-training ingredient.
- [9]
Stage 3, the teacher, went synthetic in 2023: Stanford's Alpaca showed a $600 fine-tune on GPT-generated instructions could clone much of a frontier model's behaviour, Vicuna used shared conversations and Orca used rich teacher explanations rather than bare answers.
- [10]
On-policy generalized knowledge distillation fixed the train/inference mismatch, and DeepSeek-R1 shipping a family of distilled models alongside the flagship made 'the teacher is a model' the default assumption for every small model release since.
- [11]
Stage 4, the curriculum, flipped in 2024: Self-Instruct and STaR both date to 2022, but Meta's Self-Rewarding Language Models and SPIN showed a model could generate its own tasks, judge its own outputs and improve past the ceiling of its human preference data.
- [12]
The essay describes curriculum design as historically the most artisanal part of ML, the taste-driven choice of what to train on next.
- [13]
Stage 5 is the researcher, dated 2026: the assistance era of Copilot and SWE-agents kept a human choosing the experiments, and the discovery era did not.
- [14]
DeepMind's AlphaEvolve evolved genuinely new algorithms in 2025, and Sakana's AI Scientist, now published in Nature, sketched the full paper-writing pipeline.
- [15]
Karpathy's autoresearch, in March 2026, is described as a deliberately minimal ratchet loop in which a coding agent modifies a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight.
- [16]
The stage dates given are 2022, 2023, 2023, 2024 and 2026, so two flips land in 2023 and none in 2025, and the stated one-per-year cadence does not hold across the essay's own examples.
- [17]
At the essay's stated 100x cost reduction, a fixed budget that previously funded one human-produced attempt funds about 100 simulated attempts.
- [18]
A five-minute experiment run back to back through an eight to twelve hour overnight window yields roughly 96 to 144 sequential experiments.
- [19]
The reward signal is consulted once per training sample while the corpus is assembled once per training run, so the stage that flipped first was the highest-frequency human touchpoint per unit of progress rather than the largest one-off cost.
Sources
1 independent publisher whose own reporting we read for this story.
- latent.space[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over
1 article · August 22, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Automated AI ResearchFollow
- Synthetic Training DataFollow
- Knowledge DistillationFollow
- Synthetic RL EnvironmentsFollow
- Model-Based Reward and EvaluationFollow
- Self-Improving Training LoopsFollow
Entities
- Latent SpaceFollow
- InstructGPTFollow
- Constitutional AIFollow
- MT-BenchFollow
- AlpacaEvalFollow
- Microsoft Phi seriesFollow
- WRAPFollow
- Nemotron-4 340BFollow
- AlpacaFollow
- VicunaFollow
- ORCAFollow
- DeepSeek-R1Follow
- Self-Rewarding Language ModelsFollow
- SPINFollow
- Self-InstructFollow
- STaRFollow
- AlphaEvolveFollow
- Sakana AI ScientistFollow
- autoresearchFollow
- Andrej KarpathyFollow
- Z.aiFollow
- GLM-5.3Follow
- Ornith-1.5Follow
- SimileFollow
- Generative AgentsFollow