Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

The judge went synthetic first, which tells you which part of your pipeline is next

Latent Space dates five flips from human-made to model-made since 2022. The order they happened in predicts more than the annual cadence the essay claims for it.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying The judge went synthetic first, which tells you which part of your pipeline is next
Generated illustration

What happened

  • Latent Space dates one component of the intelligence pipeline flipping from human-made to model-made each year since 2022, each with a patient-zero release at a frontier lab.
  • The corpus followed in 2023, including Apple's WRAP, which rephrased the web with an LLM and made pretraining roughly 3x more efficient.
  • Stage five is the researcher: Karpathy's March 2026 autoresearch loop edits a training setup, runs a five-minute experiment, and keeps the change only if validation loss drops.

Why it matters

  • decision It changes where an operator looks for the next substitution: the touchpoint humans are asked to service most often per iteration, not the one with the biggest invoice.
  • constraint The 100x discount only converts into quality where you can run and select among many attempts, which rules out any step with one shot and no cheap verifier.
  • exposure Once distillation is the default assumption for small model releases, shipping a strong reasoner also supplies the training signal for whoever wants to copy it.
  • contradiction The annual framing invites you to expect a flip on schedule, while the dates behind it cluster and skip, so this is a diffusion pattern rather than a clock to plan against.

The ordering is the part worth arguing about. If cost alone decided, the corpus would have gone first: text is the largest line item and the easiest thing to imitate. Instead the judge went first, in 2022, when InstructGPT collected human preferences once and let the policy optimise against a trained reward model rather than against people [3]. A reward signal is consulted once per sample. A pretraining corpus is assembled once per run. The stage that flipped earliest was the one where humans sat in the loop the most times per unit of progress [19], and once LLM-as-judge became the default evaluation method, approval and scoring both ran model-on-model [5].

The price of that substitution is stated plainly in the essay: roughly 10% worse, 100x cheaper, 10,000x faster [2]. At a hundredth of the cost, the same budget buys about a hundred attempts instead of one [17]. That arithmetic only pays where you can select or verify among the attempts. Where you get one shot and no cheap check on the output, a 10% quality penalty is just a 10% quality penalty.

Which is why the last two flips in the sequence are the interesting ones. Curriculum design, described here as historically the most artisanal and taste-driven part of ML [12], flipped in 2024 when self-rewarding models and SPIN showed a model generating its own tasks, grading its own outputs and improving past the ceiling of its human preference data [11]. Stage five, the researcher, arrives with Karpathy's autoresearch in March 2026: a coding agent edits a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight [15]. In an eight to twelve hour window that is roughly 96 to 144 sequential experiments [18].

Note what makes that loop possible. The acceptance test is validation loss, a metric that costs nothing to consult. Stage five is parasitic on stage one: the researcher can only be automated because the judge already was. That is the useful part of the framework for anyone auditing their own stack. The candidate for the next flip is not the most expensive human input, it is the human input you consult most often per iteration and can replace with something a machine can check. If your acceptance test is taste with no metric behind it, the flip stalls, which is exactly why curriculum needed a self-reward signal to arrive first.

Two cautions. This is one publisher's mental model, offered as something you arrive at if you squint at the same reading list [1] [2], and its own dates do not march to an annual beat: two flips land in 2023, none in 2025 [16]. And the essay groups end-to-end RL environments with synthetic data and synthetic rubrics as the same family of increasingly ambitious human simulation [2], while the five numbered stages we have run reward [3], data [6], teacher [9], curriculum [11] and researcher [13]. Environments get named as the pattern and then not dated. On the framework's own logic that is a stage waiting for its patient zero, not a stage that skipped its turn.

What to watch

  • Whether a ratchet-style autoresearch loop shows up inside a frontier lab's published training pipeline rather than a personal repo.
  • Whether end-to-end RL environments get a dated patient-zero release, which is the stage the framework names but does not place.
  • Whether anyone publishes the cost and quality measurements behind the 10% worse, 100x cheaper claim instead of asserting it.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence42
Adoption58
Hype gap+32
Incentives55
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Latent Space argues that every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made, and that each flip has a patient zero: a paper or product where the synthetic version first became load-bearing at a frontier lab, after which it diffuses.

    ReportedSupportedSource: latent.spaceView cited source
  2. [2]

    The essay frames synthetic data, synthetic rubrics, the AI researcher and end-to-end RL environments as one family of increasingly ambitious human simulation that is roughly 10% worse, 100x cheaper and 10,000x faster.

    ReportedSupportedSource: latent.spaceView cited source
  3. [3]

    Stage 1, the reward signal, went synthetic in 2022: InstructGPT established collecting human preferences once, training a reward model, and letting the policy optimise against the model rather than against humans.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. latent.space

    1 article · August 22, 2026

    [AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories