Security1 publisher2 min readPublished
Anthropic's Sonnet 5.5 requires a new between_tools setting for teams running thinking off
Anthropic's Claude Sonnet 5.5 keeps its listed API price and, the company says, answers more than 30% faster using fewer tokens per task. Security automation keeps less of that saving, because the model hands higher-risk cybersecurity tasks back to Sonnet 5.
The Watch · Security desk

What happened
- Developers select the model as claude-sonnet-5-5 on the Claude Platform, and it is offered with zero data retention and through AWS, Google Cloud and Microsoft Azure.
- On Terminal-Bench 4.0, an agentic test of multi-step command-line work, Sonnet 5.5 scored 70.6% against 10.3% for Sonnet 5.
- Anthropic says Sonnet 5.5 beat Sonnet 5 by 10 points on its FrontierCode test at the same High effort setting, at about one-fifteenth of the cost per task.
- An automated audit of about 1,850 scenarios found Sonnet 5.5 matched or improved on Sonnet 5 in most alignment, misuse and honesty measures, though Anthropic says its tests cannot catch every failure.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- decision Teams running thinking off have to change and test their request configuration before production traffic can move to the new model name.
- cost Work that trips the higher-risk cybersecurity safeguards runs on Sonnet 5 and pays for its token use; on Anthropic's FrontierCode run, Sonnet 5 cost about 15 times as much per task.
- exposure Tooling that moves conversations between accounts is the setup most likely to run into Anthropic's new controls against capability extraction through fake accounts, and Anthropic points those users to its guidance.
The saving Anthropic advertises is per task. The listed price has not changed, so any gain depends on the model finishing a job in fewer tokens [2]. Early testers said it understood codebases quickly and grouped tool calls into fewer steps [15]. Against Opus 5.5, Anthropic says Sonnet 5.5 costs less per task at lower effort settings and reaches comparable benchmark results at a similar cost at higher ones [8]. The cost comparisons are Anthropic's own [7][8].
Effort settings decide which comparison applies. The default is Medium in Claude Code and Anthropic's apps, and High on the Claude Platform [10]. An API integration on the Claude Platform that leaves effort unset runs at High [10].
Teams running Sonnet with thinking turned off have a precondition. They must switch to the new between_tools setting before moving to Sonnet 5.5 [4]. The report does not describe what between_tools changes in a request, or what a thinking-off call to claude-sonnet-5-5 returns without it. The edit touches two fields, in order: the thinking configuration first, then the model name [4][3].
Anthropic's cybersecurity safeguards keep routine development work on the new model, including finding and fixing bugs [11]. The older model that picks up higher-risk work is the one David Loker, VP of AI at CodeRabbit, described when he said "Sonnet 5's tendency to reach for web search too often and its high token use are both gone in this new model" [11][9].
CodeRabbit is moving in stages. "We plan to move simple and moderate reviews over now, and more in the coming weeks," Loker said [9].
Biology safeguards carry over unchanged from Sonnet 5, and Anthropic says they may flag some legitimate microbiology and virology requests [12]. Organizations can apply to Anthropic's verification programs for expanded access [12].
What to watch
- Anthropic documentation on what between_tools changes, and whether thinking-off requests to claude-sonnet-5-5 fail or degrade without it.
- A published definition of which cybersecurity tasks count as higher-risk, and whether verification programs restore Sonnet 5.5 access for security teams.
- Independent reruns of Terminal-Bench 4.0 and FrontierCode that test Anthropic's per-task cost claims.