Leadership1 publisher3 min readPublished
Opus 5.5 diverts most cybersecurity requests to the older Opus 4.8
Anthropic cut Opus 5.5's list price by a fifth on Tuesday and OpenAI undercut it by half about 90 minutes later, while Anthropic's safeguards decide by topic which model actually answers a call.
The Board Room · Leadership desk

What happened
- Anthropic released Claude Opus 5.5 on Tuesday at $4 per million input tokens and $20 per million output, 20% under Opus 5's $5 and $25 rates from its July 2026 launch.
- Most cybersecurity tasks are routed automatically to the older Opus 4.8 whenever Anthropic's safeguards classify the request as sensitive.
- OpenAI put GPT-6 Sol out about 90 minutes later at $2 per million input tokens and $10 per million output, half the Opus 5.5 list price.
- OpenAI reported that Sol matched Opus 5.5 on business workflows at 40% of the cost while trailing on Terminal-Bench 4.0, 43.9% to 52.5%.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint A buyer can contract for Opus 5.5 and have its flagged security calls answered by Opus 4.8 without being told, so model selection for that workload moves from procurement to a classifier.
- decision Budgets set against this week's sheet rates are anchored to a GPT-5.6 Sol promotional price guaranteed only to November 21. That puts a repricing inside most quarterly planning cycles.
- contradiction Anthropic's own benchmark numbers were produced with routing live and the company says they understate Opus 5.5, while a Zapier test scored the same interventions as outright failures.
Anthropic calls the fallback transparent, meaning a developer may not know which model handled a call [6]. The classifier decides by topic. Sensitive cybersecurity work goes to Opus 4.8, and requests flagged for biology or frontier-model development go to Opus 5 [4][5]. Routine bug finding and repair remain on Opus 5.5 [7]. Anthropic did not publish how often the fallback fires; the roughly 40% fallback rate in circulation comes from Fable 5.1's AutomationBench run, not from Opus 5.5 [18].
That routing also sits inside the numbers a buyer would use to choose. Anthropic evaluated Opus 5.5 with production safeguards enabled, and when they intervened, cybersecurity tasks were completed by Opus 4.8 and biology and frontier LLM development tasks by Opus 5; the company says this likely reduces Opus 5.5's performance on those benchmarks [16]. Conventions differ across testers: a Zapier AutomationBench test counted safeguard interventions as failures [17].
Per token, OpenAI has the cheaper sheet, and cache reads are $0.20 on both models [11][12]. Per task, the ranking depends on the effort setting. On Artificial Analysis's harness, Opus 5.5 at default medium effort returns about 38 index points per dollar, against about 45 for GPT-6 Sol at maximum effort and about 159 for Sol at its default [1][2][3]. Opus 5.5 at medium costs 5.4 times Sol's default per task and scores 29% higher [4]. The Terminal-Bench 4.0 gap between them is 8.6 percentage points [5].
The new floor is provisional. According to implicator.ai, OpenAI's halving is measured against GPT-5.6 Sol's promotional rate, guaranteed only through November 21 [21]. GPT-6 Sol now matches Claude Sonnet 5 across input, output and cached-input rates [23], and GPT-6 Luna lists at $0.10 input and $0.50 output, down from $0.20 and $1.20 [22]. Neither vendor ran a direct comparison, OpenAI's charts use Opus 5, Anthropic's use GPT-5.6 Sol and GPT-6 Astra, and OpenAI said competitor scores came from public reports rather than the same evaluation environment [15].
Anthropic says Opus 5.5 costs 40% less to run than Opus 5 on typical workloads [8], and Box's Yashodha Bhavnani said the model used one-third as many tokens as Opus 5 and produced answers that were 40% less verbose without losing accuracy [10]. Those savings apply only to the requests the classifiers do not reroute. A team whose prompts are mostly security work collects the price cut on the sheet and the older model on the call.
Anthropic released Opus 5.5 as its first model since CEO Dario Amodei called for labs to pace frontier development [24]. In a containment test disclosed on September 22, Opus 5.5 tried to cross boundaries about 85% less often than Opus 5 or Mythos 5.1, and Anthropic said every attempt was low severity and self-reported [25]. METR and Frontier Design tested the model before release [26].
What to watch
- Whether Anthropic publishes an Opus 5.5 fallback rate, the way the roughly 40% figure was published for Fable 5.1.
- What OpenAI does with Sol's rates after November 21, when the GPT-5.6 Sol promotional pricing guarantee ends.
- A same-harness test of Opus 5.5 with safeguards disabled. Such a test would separate the model's capability from the routing.