Build1 distinct publisher3 min readUpdated
Latent Space dates five flips from human-made to model-made since 2022. The order they happened in predicts more than the annual cadence the essay claims for it.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The ordering is the part worth arguing about. If cost alone decided, the corpus would have gone first: text is the largest line item and the easiest thing to imitate. Instead the judge went first, in 2022, when InstructGPT collected human preferences once and let the policy optimise against a trained reward model rather than against people [3]. A reward signal is consulted once per sample. A pretraining corpus is assembled once per run. The stage that flipped earliest was the one where humans sat in the loop the most times per unit of progress [19], and once LLM-as-judge became the default evaluation method, approval and scoring both ran model-on-model [5].
The price of that substitution is stated plainly in the essay: roughly 10% worse, 100x cheaper, 10,000x faster [2]. At a hundredth of the cost, the same budget buys about a hundred attempts instead of one [17]. That arithmetic only pays where you can select or verify among the attempts. Where you get one shot and no cheap check on the output, a 10% quality penalty is just a 10% quality penalty.
Which is why the last two flips in the sequence are the interesting ones. Curriculum design, described here as historically the most artisanal and taste-driven part of ML [12], flipped in 2024 when self-rewarding models and SPIN showed a model generating its own tasks, grading its own outputs and improving past the ceiling of its human preference data [11]. Stage five, the researcher, arrives with Karpathy's autoresearch in March 2026: a coding agent edits a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight [15]. In an eight to twelve hour window that is roughly 96 to 144 sequential experiments [18].
Note what makes that loop possible. The acceptance test is validation loss, a metric that costs nothing to consult. Stage five is parasitic on stage one: the researcher can only be automated because the judge already was. That is the useful part of the framework for anyone auditing their own stack. The candidate for the next flip is not the most expensive human input, it is the human input you consult most often per iteration and can replace with something a machine can check. If your acceptance test is taste with no metric behind it, the flip stalls, which is exactly why curriculum needed a self-reward signal to arrive first.
Two cautions. This is one publisher's mental model, offered as something you arrive at if you squint at the same reading list [1] [2], and its own dates do not march to an annual beat: two flips land in 2023, none in 2025 [16]. And the essay groups end-to-end RL environments with synthetic data and synthetic rubrics as the same family of increasingly ambitious human simulation [2], while the five numbered stages we have run reward [3], data [6], teacher [9], curriculum [11] and researcher [13]. Environments get named as the pattern and then not dated. On the framework's own logic that is a stage waiting for its patient zero, not a stage that skipped its turn.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Latent Space argues that every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made, and that each flip has a patient zero: a paper or product where the synthetic version first became load-bearing at a frontier lab, after which it diffuses.
The essay frames synthetic data, synthetic rubrics, the AI researcher and end-to-end RL environments as one family of increasingly ambitious human simulation that is roughly 10% worse, 100x cheaper and 10,000x faster.
Stage 1, the reward signal, went synthetic in 2022: InstructGPT established collecting human preferences once, training a reward model, and letting the policy optimise against the model rather than against humans.
Constitutional AI had the model critique itself against a set of principles (RLAIF), and Lee et al. later showed AI feedback matching human feedback at a fraction of the cost.
LLM-as-judge became the default evaluation methodology via MT-Bench and AlpacaEval, at which point reward, critique and evaluation all ran on models judging models.
Stage 2, the training data, went synthetic in 2023: Microsoft's Phi series trained a small model on LLM-synthesised textbook-quality data that punched above its parameter count, and phi-1.5 confirmed it was not a fluke.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source synthesis over widely known primaries
Everything in the cluster comes from one newsletter essay. It names many real primary artifacts per stage, which makes the historical spine checkable in principle, but none of those primaries are supplied here, several key figures (3x pretraining efficiency, 100x cheaper, 10,000x faster, 700 experiments to 20 improvements) are relayed without method detail, and the forward-looking and vendor items are unverified. The internal date contradiction is the strongest evidentiary weakness because it is visible inside the source itself.
Early stages are industry default, late stages are single-vendor
Adoption is genuinely high for the first three stages as reported: LLM-as-judge as default eval, licensed synthetic data pipelines shipped as product features, reasoning-trace corpora as standard ingredients, and distilled model families as the norm for small-model releases. Adoption thins sharply at the researcher and environment stages, where the observations are one demo run and two vendor releases in a single week, one of which is explicitly a claim.
Framing outruns the supplied evidence
The retrospective substance is solid, but the packaging is overstated in three specific ways: the annual-cadence claim is contradicted by the essay's own dates, the 10%-worse / 100x-cheaper / 10,000x-faster triplet is asserted with no derivation, and stages 5 through 7 are declared flipped on the strength of one unattended experiment run and two same-week vendor releases. The measured throughput that is available is real but modest, an 11% reduction in time-to-GPT-2, which is a smaller thing than 'the researcher went synthetic'.
Self-referential newsletter promotion of its own coverage and subjects
The essay opens by routing readers through the publisher's own reading list, its GLM coverage, its Poolside and AI-for-Science themes and the podcast episode published that day, and the concluding stage is built around Simile, a company featured in that same podcast. That is a visible incentive to present a tidy, escalating narrative and to treat vendor claims sympathetically. There is no evidence in the supplied material of paid placement or undisclosed financial interest, so this is editorial and audience incentive rather than demonstrated conflict.
Moderate on the retrospective, low on the frontier
Confidence is reasonably high that stages 1 through 4 describe real, widely adopted practice, since the named artifacts are numerous, specific and mutually corroborating within a coherent lineage. Confidence is low on the cadence framing, the multiplier figures and the 2025-2026 stages, all of which rest on one publisher's account with promotional cross-references and no second source anywhere in the cluster.
build
Z.ai pays for ZCode users in tokens, not cash: 100 million each to 50,000 signups1 distinct publisher
build
Open weights caught up on finding bugs. They did not catch up on using them.1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
build
GLM-5.3 is a paper, not an endpoint: Z.ai publishes research before weights1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026