Build1 distinct publisher3 min readUpdated
latent.space argues models and harnesses improved together, and that models keep swallowing the harness. If so, most scaffolding you write is scheduled for deletion. Permissions are not.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Absorption is not uniform, and the uneven part is where the planning happens. The harness, as latent.space defines it, is everything outside the weights that lets a model perceive through context, act through tools, persist through memory and compaction, and enforce its boundaries through permissions and guardrails [5]. The first three describe work a model could plausibly learn to do inside itself, and the route was sketched early: Toolformer, from Meta in February 2023, suggested tool use could be trained in rather than prompted [6]. Permissions are a different animal. They encode what the person holding the credentials is willing to lose, and no amount of post-training settles that on the model's behalf.
The compounding argument in that history has a second half the essay leaves on the table. latent.space's position is that a loop amplifies whatever capability the model already has rather than adding any [8], which means a harness author has two levers and owns only one of them. Raising per-step reliability is the lab's job. Cutting the number of steps is yours. Hold per-step reliability where it was and shrink a twenty-step task to five, and average success moves from roughly 36 percent to about 77 percent [1]. Push the other way and ask for 90 percent end to end across twenty steps, and the step needs to succeed 99.5 percent of the time [2]. That is a weights number. You do not reach it by wrapping retries around a step that fails one time in twenty.
Which is why the decomposition work, the part that turns twenty steps into five, is worth more than the loop it runs inside. Same for the surfaces that put a person in front of the result. The 2023 to 2024 phase that latent.space labels the retreat to human in the loop, the Cursor and Copilot period [11], produced the interfaces that are still in daily use, while the full-autonomy scaffolds of spring 2023 are gone [10]. Human attention has survived one cycle of this already. Attention is also the input that does not get cheaper when the model gets better, because a model with more capability generates more output that somebody has to accept or reject.
The schedule is not fast. ReAct put the reason-act-observe loop on paper in October 2022, as a prompting method, before anyone called it a harness [9], and it took about 38 months from there to the point where engineers noticed agents working [3]. So scaffolding written this quarter will probably earn its keep for a while. The rule is to stop paying for it twice: write it plainly, do not build abstraction layers to make it survive three model generations, and expect to delete rather than migrate.
The whole plan rests on a trajectory latent.space asserts rather than demonstrates: that models keep pulling harness functions into their weights and engineers keep deleting what got pulled [4]. If labs prioritise differently, the bet is wrong. Note the asymmetry in being wrong, though. Betting on absorption costs you some code you wrote quickly and threw away. Betting against it costs you a framework you maintained for years around a capability that shipped in a model update.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
latent.space argues the Christmas 2025 jump came from the confluence of model capability improving and the wrappers around the models maturing, with their curves of improvement crossing at the right moment.
The gap between what the harness asks of the model and what the model can deliver in practice equals the effectiveness of an agent, and the closing of that gap is what latent.space credits for the improvement in agents.
Around Christmas 2025, AI engineers noticed a change in agents: they started to work.
Lukasz Kaiser, one of the inventors of the Transformer, said on Unsupervised Learning in June: "The change last winter, last Christmas - it's a little hard to pin down. I mean, the harness changed and a little post-training changed and then new pre-trained models came... but it felt like a big jump which is not that easy to pin down what did it."
An agent harness is everything besides the model weights that makes the agent work: the environment, tools, context and guardrails. With it the model can perceive (context), act (tools), persist information (memory and compaction), and enforce its boundaries (permissions and guardrails).
Toolformer (Meta, February 2023) hinted that tool use could be trained in rather than prompted.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet's analytic essay with dated artifacts but no independent measurement
The cluster contains a single source. Its factual spine — ReAct, Toolformer, AutoGPT/BabyAGI, Cursor/Copilot, o1, Claude Code, the November 2022 ChatGPT baseline — is specific and dated, the compounding arithmetic is internally checkable, and one named external voice (Kaiser) is quoted. But the organizing thesis and the forward absorption claim are the author's own reasoning, only one quantitative outcome (Devin at ~15%) is attributed to a third party, and nothing is corroborated by a second publisher.
Real shipped artifacts across three years, but no usage or production-outcome figures
Adoption is evidenced by a chain of concrete releases and deployments the source dates — ReAct, Toolformer, AutoGPT/BabyAGI, IDE agents, o1, Claude Code — plus one third-party success-rate test. That shows the harness pattern moving through real products rather than staying a paper idea. What is missing is any usage disclosure: no user counts, no enterprise deployments, no telemetry showing the post-Christmas-2025 agents working at scale, and no evidence that models are measurably absorbing harness functions.
Mildly overstated: a confident forward thesis on hedged, single-outlet evidence
The essay is more careful than typical agent commentary — it foregrounds that the cause is hard to pin down, quotes Kaiser saying the same, and uses failure arithmetic to argue against premature autonomy. The overstatement is narrower: 'agents started to work' and the absorption forecast are presented as settled directional truths while resting on one outlet's periodization, one undated third-party benchmark, and zero production evidence. The headline implication that most scaffolding is scheduled for deletion outruns what the supplied material can demonstrate.
No disclosures or relationships in the supplied material
The cluster provides only the publisher name, URL and publication date. There is no disclosure of commercial relationships with the vendors discussed (Anthropic, Cursor, GitHub, Cognition, Meta), no sponsorship statement, no product the author sells, and no funding information. Inferring an incentive structure from an outlet's beat alone would be speculation, so this dimension is left unmeasured.
Low-to-moderate: internally coherent argument, structurally unverified
Confidence is limited by single-publisher sourcing and by the mix of claim kinds. Dated historical facts and the arithmetic are high-confidence; the causal thesis, the Christmas 2025 inflection and the absorption forecast are not independently checkable within the cluster, and one supporting figure (Devin ~15%) is cited without a date. No contradicting source is present either, so the assessment is bounded by absence of triangulation rather than by conflict.
invest
Northzone says the AI notetaker is dead. Its own portfolio is the tell1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
product
Engineering counts merged pull requests and nothing for the hours spent watching the agent1 distinct publisher
build
Z.ai pays for ZCode users in tokens, not cash: 100 million each to 50,000 signups1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026