Skip to content

Build1 publisher3 min readPublished

Four model releases in 12 days break code written against earlier versions

Google, Anthropic and OpenAI shipped four model releases in 12 days, and each one's own docs list ways that code written for the earlier version now fails. A swap of the model ID is a dependency upgrade and needs contract tests at the provider boundary before it ships.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Four model releases in 12 days break code written against earlier versions
Generated illustration

What happened

  • Release notes for Claude Sonnet 5.5, which followed on September 28, list five ways code written for Sonnet 5 can break on the new model.
  • On September 17, Google's Antigravity Agent 09-2026 moved built-in tools to PascalCase parameters and replaced write_file with write_to_file or replace_file_content.
  • GPT-6.1 Sol, released by OpenAI at DevDay on September 29, does not support the none and minimal reasoning efforts, according to its model page.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Teams still on antigravity-preview-05-2026 had 18 days between the replacement's release and the shutdown to move their tool parsing to the new names.
  • exposure A local tool runner that dispatches on the old write_file name will receive calls it does not recognise once it points at the new Antigravity agent.
  • decision OpenAI clients that send the none or minimal reasoning effort have to pick a supported level before their model string can move to GPT-6.1 Sol.
  • precedent With three vendors shipping documented breaks inside 12 days, a contract-test gate becomes a recurring cost of every model swap, budgeted like any dependency bump.

Editing `model: "claude-sonnet-5"` looks like a config change. A dev.to tutorial that collected September's breaks argues the string "is part of an API contract" [20]. Its author wrote: "Changing a model string is a dependency upgrade. It deserves a test suite." [13]

The loud Sonnet 5.5 breaks fail on the first call. A client that turns off up-front thinking with `"disabled"` now has to send `thinking: {"type": "between_tools"}`, at high effort or below [7]. Forced tool use through `tool_choice` types `any` and `tool` returns a 400 error here too [8]. On the Claude API and Google Cloud, the older `computer_20251124` computer use tool is not accepted [10]. The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors [11]. Each of these fails the first time a test sends it. A 400 at least has the decency to show up in the error logs.

The quiet break is the one I would test first. Anthropic's What's new page says text between tool calls now comes back in `thinking` blocks, a change that "alters the response shape without failing any request" [12]. An agent UI that shows that narration as a progress line reads `text` blocks. After the swap it gets a 200 and has nothing to render. Anthropic deserves credit for writing down a change that no error would have exposed. The tutorial's mock reproduces it: Sonnet 5 returns the note as a text block, and the Sonnet 5.5 mock returns a thinking block that stays empty unless the request asked for `between_tools` [17]. The author says the mock's request rules follow the docs and its replies are made up [14]. The empty string is a model of the failure, not a captured API response.

The gate has two parts. One check encodes the documented breaks as rules inside a mock provider. The other runs the same contracts against two model versions and diffs the results [15]. Both sit behind a request and response type the app owns. "Your contracts should describe your app, not the provider," the tutorial says [16]. The mock's `tool_choice` error string is copied from Anthropic's docs [19]. The whole thing runs on Node.js 18 or newer with TypeScript and tsx, and it needs no API key [18][14].

The known-breaks check is only as complete as the page it was built from. Anthropic's release notes number five breaks, and the response-shape change sits on a second page, so Sonnet 5.5 has six documented changes across two pages [4]. A gate built from the release notes alone would pass a build that drops every progress message. The diff check closes that gap only when it runs against real endpoints, where replies vary from run to run. For a passing mock run to mean anything in production, the diff has to assert on structure: block types, stop reasons, tool names. I would not diff text at all.

One more test belongs in the suite. The notes say "Thinking blocks are tied to the model and the conversation" [9]. A team that stores conversations with thinking blocks and replays them after the swap should run a stored Sonnet 5 transcript against 5.5 before trusting it.

What to watch

  • Whether Anthropic's next release notes fold response-shape changes like the thinking-block move into the numbered list of breaks.
  • Whether Google holds the October 5, 2026 shutdown date for antigravity-preview-05-2026.
  • A run of the tutorial's two-version diff against live Sonnet 5 and Sonnet 5.5 endpoints, to show whether structural assertions stay stable across real, varying replies.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories