Build1 publisher3 min readPublished
Simon Willison credits two November model releases with making coding agents reliable for daily use
Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Willison called both releases incremental improvements on earlier models, and he found the change in how each one performed inside its own coding agent harness.
- By his count, about 40 of the 277 sessions at WeAreDevelopers World Congress North America touched on sandboxing or agent security.
- Willison had predicted hijacked coding agents would cause real-world economic damage this year, and he wrote that it has not really happened.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams adopting agents have to evaluate the model and the harness as one unit, because Willison credits the reliability to the pairing. Swapping either part means testing again.
- constraint General model benchmarks are a weak signal for agent adoption. The threshold Willison describes showed up only in sustained use, so teams need trials on their own tasks.
- exposure Putting agents into daily work gives them more to reach before sandboxing is solved. Willison predicted that problem would be fixed this year, and it still takes up a large part of conference programs.
The unit Willison describes is a pairing. He wrote that the two models, when paired with their respective coding agent harnesses, went from often making mistakes to being reliable enough to use day to day [5]. The harnesses were older than the models. Claude Code had been around since February 2025, and Codex was a little younger [6]. That leaves about nine months between Claude Code's arrival and the November model releases that, in Willison's account, made it dependable [2] [7].
He described the models themselves as modest steps. They were, he wrote, "incremental improvements on the models that came before them" [3]. His explanation is a threshold: "every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working" [4]. His own test did not pick up the line. He asks every new model to generate an SVG of a pelican riding a bicycle, and calls it "probably the world's stupidest benchmark" [8]. His November verdict: "Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too." [9] A bird that cannot ride a bicycle was never going to say whether a migration script runs clean. The capability showed up through use, as developers tinkered with the new model and agent combinations over the December holidays [10].
For another team, the evidence is one engineer's year of daily work. The notes do not include an error rate or a description of the tasks behind the judgement. For his result to transfer, a team would need to run the same model and harness pairs, on work that resembles his, with a bar for "reliable" set where he sets it. I think that is enough to justify a trial on real tickets and too little to build a staffing plan on.
Willison's own response was to push. He reversed his usual resolution to take on fewer projects: "We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!" [11] He wrote that "the only way to find the limits of this technology is to keep on pushing them until they don't work" [12].
On code quality he treats one of his forecasts as settled. Of his prediction that "it will become undeniable that LLMs write good code," he wrote: "I think we're there now." [13] Security is further behind. He had predicted that sandboxing would finally be solved [18]. At WeAreDevelopers World Congress North America in San Jose, where he gave the closing keynote [1], he counted around 40 of 277 sessions that touched on sandboxing or agent security [14]. The share is about 14% [15]. The "Challenger disaster" he forecast, with coding agents hijacked and causing real-world economic damage, "hasn't really played out," he wrote [16].
His summary of the year was personal. "As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically," he wrote [17].
What to watch
- A published error rate or task-level evaluation for the Claude Opus 4.5 and GPT-5.1 agent pairings, which would show whether Willison's threshold holds beyond his own work.
- The first documented case of a hijacked coding agent causing economic damage, the disaster Willison predicted and says has not yet happened.
- Willison's year-end verdict on his resolution to take on as many new projects as he likes.