Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Write your agent harness like a lease: most of it has an absorption date

latent.space argues models and harnesses improved together, and that models keep swallowing the harness. If so, most scaffolding you write is scheduled for deletion. Permissions are not.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Write your agent harness like a lease: most of it has an absorption date
Generated illustration

What happened

  • latent.space dates the moment agents began to work reliably to around Christmas 2025, when engineers paired the newest harnesses with the newest models.
  • Lukasz Kaiser, one of the Transformer's inventors, said the jump was hard to attribute: the harness changed, post-training changed, and new pre-trained models arrived.
  • The essay's frame: agent effectiveness is the gap between what a harness asks of a model and what the model can deliver, and that gap closed.
  • Its widest-gap case is spring 2023, when AutoGPT and BabyAGI handed full autonomy to models that could not carry it.
  • The prediction that follows: weights keep absorbing harness functions, engineers keep deleting the absorbed parts, and what survives is scaffolding aimed at human attention.

Why it matters

  • cost Harness code with an absorption date belongs in the maintenance budget as a write-off, not in the asset column. The team that abstracts it for longevity pays for it twice.
  • constraint Loop design cannot buy task length. Until per-step reliability moves, long autonomous tasks stay out of reach no matter how good the orchestration code is.
  • decision Next quarter's engineering has to be allocated between reasoning scaffolds a lab may ship for free and the permission and review surfaces nobody will ship for you.
  • precedent Any prompting trick can now be scoped on the assumption that it gets trained in later, which changes how much structure is worth building around it.

Absorption is not uniform, and the uneven part is where the planning happens. The harness, as latent.space defines it, is everything outside the weights that lets a model perceive through context, act through tools, persist through memory and compaction, and enforce its boundaries through permissions and guardrails [5]. The first three describe work a model could plausibly learn to do inside itself, and the route was sketched early: Toolformer, from Meta in February 2023, suggested tool use could be trained in rather than prompted [6]. Permissions are a different animal. They encode what the person holding the credentials is willing to lose, and no amount of post-training settles that on the model's behalf.

The compounding argument in that history has a second half the essay leaves on the table. latent.space's position is that a loop amplifies whatever capability the model already has rather than adding any [8], which means a harness author has two levers and owns only one of them. Raising per-step reliability is the lab's job. Cutting the number of steps is yours. Hold per-step reliability where it was and shrink a twenty-step task to five, and average success moves from roughly 36 percent to about 77 percent [13]. Push the other way and ask for 90 percent end to end across twenty steps, and the step needs to succeed 99.5 percent of the time [14]. That is a weights number. You do not reach it by wrapping retries around a step that fails one time in twenty.

Which is why the decomposition work, the part that turns twenty steps into five, is worth more than the loop it runs inside. Same for the surfaces that put a person in front of the result. The 2023 to 2024 phase that latent.space labels the retreat to human in the loop, the Cursor and Copilot period [11], produced the interfaces that are still in daily use, while the full-autonomy scaffolds of spring 2023 are gone [10]. Human attention has survived one cycle of this already. Attention is also the input that does not get cheaper when the model gets better, because a model with more capability generates more output that somebody has to accept or reject.

The schedule is not fast. ReAct put the reason-act-observe loop on paper in October 2022, as a prompting method, before anyone called it a harness [9], and it took about 38 months from there to the point where engineers noticed agents working [15]. So scaffolding written this quarter will probably earn its keep for a while. The rule is to stop paying for it twice: write it plainly, do not build abstraction layers to make it survive three model generations, and expect to delete rather than migrate.

The whole plan rests on a trajectory latent.space asserts rather than demonstrates: that models keep pulling harness functions into their weights and engineers keep deleting what got pulled [16]. If labs prioritise differently, the bet is wrong. Note the asymmetry in being wrong, though. Betting on absorption costs you some code you wrote quickly and threw away. Betting against it costs you a framework you maintained for years around a capability that shipped in a model update.

What to watch

  • Whether the next round of model releases ships trained-in memory and compaction, which would date a large amount of currently hand-written context plumbing.
  • Whether any lab publishes per-step reliability figures for agentic tasks, so step budgets can be computed rather than guessed at.
  • Whether a cleaner causal account of the Christmas 2025 jump emerges, separating post-training gains from harness maturity.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence42
Adoption48
Hype gap+18
Incentives
Insufficient
Confidence38
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    latent.space argues the Christmas 2025 jump came from the confluence of model capability improving and the wrappers around the models maturing, with their curves of improvement crossing at the right moment.

  2. [2]

    The gap between what the harness asks of the model and what the model can deliver in practice equals the effectiveness of an agent, and the closing of that gap is what latent.space credits for the improvement in agents.

  3. [3]

    Around Christmas 2025, AI engineers noticed a change in agents: they started to work.

    ReportedSupportedSource: latent.spaceView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. latent.space

    1 article · August 22, 2026

    The Evolution of the Agent Harness

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories