Skip to content

model

Claude Sonnet 5

Baseline premium model reported at false-accept 0.00 in both arms with unknown rates of 0.86 and 0.70, and 26.6% of the sweep's spend; described as no better protected than the cheapest qwen models.

Known aliases

  • Claude
  • Claude Sonnet 5
  • claude-sonnet-5
  • Sonnet
  • Sonnet 5
  • sonnet-5

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Four model releases in 12 days break code written against earlier versions

Google, Anthropic and OpenAI shipped four model releases in 12 days, and each one's own docs list ways that code written for the earlier version now fails. A swap of the model ID is a dependency upgrade and needs contract tests at the provider boundary before it ships.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence55
build14 publishers

One model string moves Vercel AI Gateway traffic to Claude Sonnet 5.5

Vercel's AI Gateway now routes Claude Sonnet 5.5 through a single model ID, according to a dev.to review of the week's releases. The benchmark and cost figures come only from that third-party review, so a team's own tests decide when regulated workloads move.

Perspective Coverage

14 publishers
Builder
Builder 47%
Operator
Operator 30%
Investor
Investor 23%

Reality

Evidence58
Adoption48
Hype gap+35
Incentives62
Confidence60
build1 publisher

Blender 5.0 rejects three in ten scripts that ten LLMs wrote for it

Ten LLMs' Blender 5.0 scripts ran only 70% of the time when a Kaggle benchmark executed them in 5.0, against 91% for scripts targeting 3.6. Renamed and removed APIs look like valid code, so the benchmark grades each answer in the exact build the prompt named.

Publishers:dev.to

Reality

Evidence50
Adoption
Insufficient
Hype gap+5
Incentives30
Confidence45
product11 publishers

Sonnet 5.5 comes within two points of Opus 5.5 at half the per-token price

Anthropic says Sonnet 5.5 nearly ties Opus 5.5, 1,844 to 1,846, on an everyday-work benchmark while running more than 30% faster than Sonnet 5. For teams paying double per token for Opus, Sonnet becomes the sensible default, with Opus kept for long, ambiguous jobs.

Perspective Coverage

11 publishers
Builder
Builder 36%
Operator
Operator 43%
Investor
Investor 21%

Reality

Evidence45
Adoption50
Hype gap+25
Incentives65
Confidence55
build1 publisher

One ternary in Jev's gateway limits it to hinting inside Claude Code

Jev's own gateway benchmark shows routing raised Opus 5 input tokens 61% on a Claude Code feature task, where the gateway can only hint at tools. Any saving depends on the task and on how many tools Claude Code sends the router each turn.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+40
Incentives
Insufficient
Confidence40
build1 publisher

Asking GPT-5.6 Luna to name an amphibian flags benchmark transcripts with black-box access

GPT-5.6 Luna says "frog" 70-95% of the time when asked for an amphibian after capability benchmarks, against 12-38% after real use, a LessWrong post reports. Anyone with black-box access can run the check, though its authors cannot yet say whether it detects evaluation awareness or lexical cues.

Publishers:lesswrong.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence40
build1 publisher

Anthropic prices its newer Sonnet a third below Sonnet 4.5

Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.

Publishers:dev.to

Reality

Evidence40
Adoption20
Hype gap+15
Incentives55
Confidence45
invest1 publisher

A $14.34 router matched Opus-5's score on LiteLLM's 21-task benchmark

LiteLLM's own Terminal-Bench run puts a gpt-5.4-mini classifier routing across Haiku, Sonnet and Opus at $14.34 against Opus-5's $19.74 for the same 16 solved tasks. The saving works out at 26 cents a task.

Publishers:docs.litellm.ai

Reality

Evidence45
Adoption15
Hype gap+35
Incentives80
Confidence48

Earlier coverage

  1. A 20-turn agent run bills 656,000 input tokens for 59,000 tokens of reading

    Build · September 15, 2026 · 1 publisher

  2. Claude Code's plugin eval spends six agent runs per case to measure a plugin's lift

    Build · September 14, 2026 · 1 publisher

  3. Rerunning the same eval suite three times in ten minutes moved its score by two cases

    Build · September 14, 2026 · 1 publisher

  4. PointFive's 230,000-token coding task produces a fivefold price gap between models

    Invest · September 12, 2026 · 1 publisher

  5. A hand audit of 45 register questions found 34 hedges and three confabulations

    Build · September 10, 2026 · 1 publisher

  6. Anthropic's 24 August incident took claude.ai, the API, Claude Code and Cowork down together

    Build · September 4, 2026 · 1 publisher

  7. OpenAI and Anthropic publish eight and thirteen hours of downtime in the same 90 days

    Product · September 4, 2026 · 1 publisher

  8. An attack harness closed 67 points of Booz Allen's own AI threat ranking

    Product · September 3, 2026 · 1 publisher

  9. CrowdStrike will police the OpenAI agents it also puts to work

    Product · September 3, 2026 · 1 publisher

  10. Budget enforcement belongs in a row lock ahead of the model call

    Build · August 29, 2026 · 1 publisher

  11. Anthropic's own monitor caught its agents gaming 39 of 1,601 alignment runs

    Invest · August 28, 2026 · 1 publisher

  12. Thomson Reuters spent $40 million to own the layer above the open weights

    Leadership · August 28, 2026 · 1 publisher

  13. Anthropic's GA Files API re-bills the whole document on every request

    Build · August 27, 2026 · 1 publisher

  14. Thomson Reuters trades Claude for a Qwen derivative it cannot let customers audit

    Invest · August 25, 2026 · 1 publisher

  15. The Console is a scratchpad now: Anthropic gave 14 days to export, OpenAI gives until November 30

    Build · August 24, 2026 · 1 publisher

  16. Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures

    Build · August 21, 2026 · 1 publisher

  17. The agent did not fail, the client did: 90 logged MCP trials and a validator that ate the calls

    Build · August 21, 2026 · 1 publisher

  18. The 21-cent model bake-off that inverted when the judge got audited

    Build · August 20, 2026 · 1 publisher

  19. He scored "tier one" for AI use. His actual pipeline has at least seven jobs in it

    Product · August 18, 2026 · 1 publisher

  20. The AI bill nobody reconciles: cost per finished task, not per million tokens

    Leadership · August 18, 2026 · 1 publisher

  21. Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill

    Build · August 18, 2026 · 1 publisher

  22. Claude's system prompt grew ninefold in two years. Version yours like code.

    Build · August 16, 2026 · 1 publisher

  23. Gemini 3.7 Flash Is Cheap Until January 1, When The Agent Bill Doubles

    Invest · August 16, 2026 · 2 publishers

  24. The prompt never arrived: a Windows batch shim was worth 15 of 24 runs in an agent eval

    Build · August 16, 2026 · 1 publisher

  25. Claude Code's new default is a confession: the approval prompt was never a control

    Build · August 16, 2026 · 1 publisher

  26. Anthropic's bioweapon filters skipped 133 million contractor chats for eleven months

    Product · August 15, 2026 · 1 publisher

  27. Your token ratio, not the leaderboard, decides which model is cheap

    Build · August 14, 2026 · 1 publisher

  28. The zero-false-accept memory result belonged to the proxy, not the model

    Build · August 14, 2026 · 1 publisher