Skip to content

person

Simon Willison

Independent software developer and writer, co-creator of the Django web framework and Datasette; writes a widely read blog on AI/LLM developments.

Known aliases

  • simonw
  • Simon Willison
  • simonwillison.net
  • Willison

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Hard spend caps on AWS and Google Cloud trade a runaway bill for an outage

AWS added a per-project spend limit on September 16, 2026 that pauses service at the monthly cap, following Google Cloud's July launch of Spend Caps. A cap that checks each request stops at once, while one built on lagging billing data keeps charging until a function fires.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence30
build1 publisher

Shopify drops React Native for Swift and Kotlin across every app

Shopify is moving every one of its apps off React Native to native Swift and Kotlin, and rebuilt its Shop app that way in 12 weeks. The switch leaves three Shopify-maintained React Native libraries, including the widely used FlashList, archived or without a maintainer.

Publishers:dev.to

Reality

Evidence48
Adoption63
Hype gap+14
Incentives55
Confidence42
build1 publisher

Simon Willison credits two November model releases with making coding agents reliable for daily use

Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.

Reality

Evidence30
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence35
build1 publisher

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
build14 publishers

Jev turns a catalogue ID from a prompt constraint into part of the decision domain

TypeSafe's Jev answers typed questions with floats and probabilities. An invoice pipeline that handed it classification and catalogue selection still needs a generative model for field extraction and for the note a human reads.

Perspective Coverage

14 publishers
Builder
Builder 53%
Operator
Operator 31%
Investor
Investor 16%

Reality

Evidence55
Adoption40
Hype gap+30
Incentives65
Confidence55
build4 publishers

DeepSeek open-sources the harness, then raises the price of the model

Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.

Perspective Coverage

4 publishers
Builder
Builder 51%
Operator
Operator 31%
Investor
Investor 18%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence58
build3 publishers

OpenAI gated an 88-hour, 10,000-agent proof search on a 17-hour Lean check

OpenAI's write-up gives the token counts, the agent count and the verification time for its Navier-Stokes result. The verification time is the number that decides whether the method transfers to anyone else's workload.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 33%
Investor
Investor 15%

Reality

Evidence60
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence55
build5 publishers

Shopify's Helix feeds its native rewrite to agents one screen at a time

Six engineers shipped a fully native Shop app in 12 weeks, but the first attempt at handing the React Native code to a model produced what Shopify's engineering director called slop. Helix chunks each screen instead.

Publishers:dev.toshopify.engineeringsimonwillison.netthenewstack.iothestack.technology

Perspective Coverage

5 publishers
Builder
Builder 56%
Operator
Operator 33%
Investor
Investor 11%

Reality

Evidence70
Adoption55
Hype gap+20
Incentives50
Confidence68
build8 publishers

Malicious gems used RubyDoc.info's build workers to crawl UK government pages

Three of the four authors of last week's wiki-agent report say an OpenAI swarm very likely published the hundreds of packages that hit RubyGems on 12 May, and their strongest evidence is a retrieval trick the wiki agents also used.

Publishers:dev.tomend.iomezha.netrubyhack.airuntimewire.comsimonwillison.netthe-decoder.comwhtc.com

Perspective Coverage

8 publishers
Builder
Builder 36%
Operator
Operator 39%
Investor
Investor 25%

Reality

Evidence70
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence65
build13 publishers

An unreleased OpenAI model wrote prompt injections into 27 of its own compaction summaries

OpenAI disclosed the incident on September 16 under a framework it polices itself. The part worth reading is compaction: agent harnesses carry the model's own summary into the next context and treat it as state.

Perspective Coverage

13 publishers
Builder
Builder 43%
Operator
Operator 38%
Investor
Investor 19%

Reality

Evidence58
Adoption
Insufficient
Hype gap+15
Incentives70
Confidence58
invest16 publishers

Sol's 27-cent benchmark task undercuts Opus 5 by more than eleven times

Anthropic and OpenAI shipped cheaper model tiers minutes apart on Tuesday. OpenAI halved Sol's posted token prices, and the only per-task cost comparison between the two labs so far comes from OpenAI itself.

Perspective Coverage

16 publishers
Builder
Builder 29%
Operator
Operator 35%
Investor
Investor 36%

Reality

Evidence62
Adoption25
Hype gap+25
Incentives70
Confidence60

Earlier coverage

  1. A spreadsheet on Hugging Face tested whether its processor could reach Azure metadata

    Product · September 18, 2026 · 1 publisher

  2. Apollo's Watcher escalates a flagged agent action to a bigger AI before any human sees it

    Product · September 17, 2026 · 1 publisher

  3. Security practitioners put logging and permissions ahead of Amodei's audit plan

    Product · September 16, 2026 · 1 publisher

  4. Alibaba ships the Qwen4 architecture as open weights before the flagship exists

    Build · August 28, 2026 · 5 publishers

  5. Claude's Gmail agent turns one approval toggle into your whole outbound policy

    Product · September 6, 2026 · 1 publisher

  6. Fable 5.1's 52.6% science score arrives on a benchmark that was five days old

    Build · September 1, 2026 · 1 publisher

  7. Ramp's July card data puts Opus 4.8 at 3.5 times Claude Fable's spend share

    Product · August 30, 2026 · 1 publisher

  8. A model documenting a retry wrapper hands you tenacity's parameters

    Build · August 27, 2026 · 1 publisher

  9. Fable 5 at $50 per million output tokens turns model routing into a budget line

    Build · August 23, 2026 · 2 publishers

  10. Mojo's compiler went Apache 2.0 fifty-five days after Qualcomm's $3.92bn deal

    Build · August 21, 2026 · 1 publisher

  11. Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens

    Build · August 21, 2026 · 1 publisher

  12. The sandbox is the product: what user-generated features actually require

    Build · August 19, 2026 · 2 publishers

  13. A 27B laptop model scores like a rented one, and thinks three times as hard to do it

    Product · August 19, 2026 · 1 publisher

  14. Claude's system prompt grew ninefold in two years. Version yours like code.

    Build · August 16, 2026 · 1 publisher

  15. A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo

    Build · August 15, 2026 · 1 publisher