Skip to content

Build1 publisher3 min readPublished

Thompson credits the harness for the agent jump that retired his bubble call

Ben Thompson now says he does not think AI is a bubble, and the evidence he offers is a capability jump that showed up in Claude Code in December, weeks after Anthropic shipped the Opus 4.5 weights.

The Engineer · Build desk

Illustration accompanying Thompson credits the harness for the agent jump that retired his bubble call

What happened

  • He frames the reversal around three LLM inflection points: the November 2022 ChatGPT launch, OpenAI's o1 in September 2024, and the agent capability that appeared late in 2025.
  • Thompson writes that agentic workloads depend on more than the model, and names the harness, the software that actually controls the model, as a critical component.
  • By his account Claude and Codex were completing tasks that ran for hours and completing them correctly, which is what he treats as the third inflection.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability If the December jump came from the harness rather than the weights, a team can add agent capability on its own release schedule instead of waiting for the next frontier model.
  • constraint As published, the argument goes as far as a capability claim, so anyone using it to defend a compute budget has to supply the utilization and token figures themselves.
  • cost Tasks that run for hours consume tokens for hours, and the bill lands on whoever owns the inference contract for the harness, not on the model vendor.
  • contradiction Thompson undercuts his own verdict in the same sentence, calling it possibly the truest evidence of a bubble. That hedge makes it weak material to quote as settled.

The date is the interesting part, and Thompson flags it himself. Anthropic released Opus 4.5 on November 24, 2025, to what he describes as relatively little fanfare [12]. The capability he is pointing at arrived later. At some point in December, he writes, Claude Code with Opus 4.5 "suddenly seemed to be able to do things that were never possible previously" [13]. Thompson credits the harness. He writes that a critical component of making agentic workloads work is the harness, the software that actually controls the model [16].

That distinction has a practical edge for anyone buying capacity. A harness is code in a repository, and what goes into the context window can be changed without a training run. OpenAI shipped GPT-5.2-Codex on December 18, twenty-four days after Opus 4.5, and Thompson says it was similarly capable [14][18]. By his account both were finishing tasks that took hours, and finishing them correctly [15].

Set against the earlier inflections, the spacing is uneven. Transformer-based LLMs were introduced in 2017 and ChatGPT launched in November 2022, roughly five years later [6][5][21]. o1 followed in September 2024, about 22 months after ChatGPT, and Opus 4.5 came about 14 months after o1 [9][19].

Each step changed who carries the error-handling. Thompson's complaint about the first ChatGPT releases was that they hallucinated, and that a user had to know what to use them for and verify the output [8]. On o1 he wrote in an Update at the time: "The big challenge for traditional LLMs is that they are path-dependent; while they can consider the puzzle as a whole, as soon as they commit to a particular guess they are locked in, and doomed to failure" [10]. Reasoning moved that self-check inside the model, and the harness moves the loop that runs the task outside it [9][16].

What the published excerpt supports is a claim about capability. It breaks off mid-sentence in the passage on Claude Code and Codex abstracting the user away from the model, and it does not include figures for capital spending, utilization or token volumes [20][17]. So a buyer who wants to cite this in a capex review is citing three model release dates and one writer's December. For the conclusion to transfer, the hours-long tasks would have to be ones that a buyer's own reviewers accept without redoing the work, and the tokens burned on them would have to be billed to a customer who renews. Thompson had previously argued that bubbles can be good [1]. Writing in March 2026, on the morning of Nvidia's GTC, he reached a different conclusion and hedged it in the same breath: "I don't think we're in a bubble (which, paradoxically, maybe is the truest evidence we are)" [2][3].

What to watch

  • Whether the rest of Thompson's essay ties agent token volume to a specific spending or utilization figure.
  • Whether the next Claude Code or Codex capability jump ships with new weights or with a harness update. That is the test of his explanation.
  • What Nvidia says at the GTC Thompson wrote on the morning of about who is consuming inference capacity.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories