Skip to content

model

Claude Opus 5

Claude Opus 5 is Anthropic's flagship large language model, built for complex reasoning and agentic coding tasks, offered in multiple speed/cost tiers.

Known aliases

  • Claude Opus 5
  • claude-opus-5
  • claude-opus-5[1m]
  • Claude Opus 5.5
  • Claude Opus 5 Fast
  • Claude Opus 5 Free Desktop
  • Claude Opus 5 High
  • Claude Opus 5 (max)
  • Claude Opus 5 (med)
  • Claude Opus 5 (xhigh)
  • Opus 5
  • Opus-5
  • Opus 5 High

Relationships

No evidence-backed relationships are recorded.

Current stories

build2 publishers

Two-thirds of failed agent runs in ThinkingBox exited cleanly with the backend still wrong

Microsoft and Hugging Face's ThinkingBox found that 67.24% of 79,853 failed agent runs ended cleanly, with no final tool error. Those failures showed up only when executable checks read the records each run left in the backend.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 38%
Investor
Investor 10%

Reality

Evidence60
Adoption
Insufficient
Hype gap+10
Incentives35
Confidence68
build1 publisher

Cantina's open-weights exploit model ran a 60-task security eval for $2.38

Cantina released apex-flash-1, an open-weights vulnerability-research model it says solved 40 of 60 tasks for $2.38, against $74.68 for Claude Opus 5 High. There is no hosted endpoint, so teams download the 321-billion-parameter weights, pay for their own inference and verify the numbers themselves.

Publishers:runtimewire.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence40
build3 publishers

Coding agents routed more than 13,000 private screenshots through public GitHub repos

Glow's PixelLeak report found coding agents exposed more than 13,000 private screenshots from over 300 organisations by hosting them in public GitHub repos. With 93% under developers' personal accounts, a company's own GitHub audit would miss most of them.

Perspective Coverage

3 publishers
Builder
Builder 42%
Operator
Operator 50%
Investor
Investor 8%

Reality

Evidence55
Adoption50
Hype gap+10
Incentives55
Confidence60
build1 publisher

Cache expiry decides whether a whole knowledge base in the prompt undercuts retrieval

Cache-augmented generation costs about what retrieval does when the corpus is roughly 10 times the tokens retrieval would send, a dev.to analysis finds. Sparse traffic breaks the rule, because each query then pays the cache-write premium and caching becomes the most expensive option.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence40
build17 publishers

Opus 5.5 matched Opus 5's puzzle answers for up to 69 percent less at its lower default effort

Opus 5.5 matched Opus 5 on two reasoning puzzles in The New Stack's tests at 43 to 69 percent lower cost. Both ran at default effort, medium on the new model and high on the old, so the saving a team sees depends on the effort level it pins.

Perspective Coverage

17 publishers
Builder
Builder 43%
Operator
Operator 33%
Investor
Investor 24%

Reality

Evidence62
Adoption48
Hype gap+22
Incentives58
Confidence58
build1 publisher

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
build1 publisher

One ternary in Jev's gateway limits it to hinting inside Claude Code

Jev's own gateway benchmark shows routing raised Opus 5 input tokens 61% on a Claude Code feature task, where the gateway can only hint at tools. Any saving depends on the task and on how many tools Claude Code sends the router each turn.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+40
Incentives
Insufficient
Confidence40

Earlier coverage

  1. DeepSeek's vision agent wins three of eleven benchmarks, all against May's Opus

    Product · August 21, 2026 · 2 publishers

  2. Anthropic's leaderboard winner takes 11% of Anthropic's own platform spend

    Build · August 24, 2026 · 2 publishers

  3. Nvidia's SoL-Pi rewrites coding-agent harnesses to use up to 49 percent fewer tokens

    Build · September 26, 2026 · 1 publisher

  4. Opus 5.5 cost less per unit of coding work than Sonnet 5, with fewer review rounds needed

    Build · September 25, 2026 · 1 publisher

  5. Runway's Solaris turns every click into conditioning data for the next generated frame

    Build · August 31, 2026 · 5 publishers

  6. Anthropic Cuts Cache-Read Prices by 75%; Cache Reads Were ~60% of a Heavy Agent's Bill Before the Cut

    Invest · September 1, 2026 · 2 publishers

  7. Meta keeps Muse Spark 1.3 pricing flat while claiming coding edge over GPT-5.6

    Product · September 3, 2026 · 3 publishers

  8. Gemini 3.8 Flash's introductory price doubles on December 31, 2026

    Build · September 2, 2026 · 8 publishers

  9. Two harnesses put the same model 37 points apart on ARC-AGI-3

    Science · September 3, 2026 · 2 publishers

  10. GitHub bills HydraFusion by every model leg its router decides to call

    Build · September 4, 2026 · 2 publishers

  11. Hacktron chained a Claude-written libheif exploit into OpenAI's internal repositories

    Security · September 18, 2026 · 9 publishers

  12. Opus 5.5 diverts most cybersecurity requests to the older Opus 4.8

    Leadership · September 22, 2026 · 2 publishers

  13. A tester left Claude Opus 5.5 running unattended for 18 hours across six repositories

    Security · September 23, 2026 · 3 publishers

  14. OpenAI halves the API price of Sol and Luna against GPT-5.6's promotional rates

    Product · September 22, 2026 · 8 publishers

  15. Box measured Claude Opus 5.5 using a third of the tokens Opus 5 needed

    Product · September 22, 2026 · 7 publishers

  16. Sol's 27-cent benchmark task undercuts Opus 5 by more than eleven times

    Invest · September 22, 2026 · 16 publishers

  17. A flagged retrieval can drop Opus 5.5 to Opus 4.8 for the rest of the conversation

    Leadership · September 23, 2026 · 1 publisher

  18. Opus 5.5's claimed 40% cost cut needs a cache-heavy workload to appear

    Science · September 23, 2026 · 2 publishers

  19. Anthropic lists pasted-prompt injection as a regression in Opus 5.5's own audit

    Security · September 23, 2026 · 1 publisher

  20. OpenAI's 50 percent API price cut doubles the token volume a flat budget buys

    Security · September 23, 2026 · 1 publisher

  21. Anthropic now ships a frontier model every 26 days, mostly by repricing the last one

    Product · September 23, 2026 · 1 publisher

  22. Anthropic cuts Opus 5.5 prices 20% on tokens, 60% on cache reads, citing fewer tokens burned for 40% total savings

    Invest · September 23, 2026 · 1 publisher

  23. Each plan-mode toggle under opusplan invalidates Claude Code's prompt cache

    Build · September 23, 2026 · 1 publisher

  24. OpenAI cuts prices on new GPT-6 Sol and Luna models

    Product · September 23, 2026 · 1 publisher

  25. A 24,000-character CJK tool result reached Claude Code's context as 49,964 tokens

    Build · September 22, 2026 · 1 publisher

  26. DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September

    Build · September 22, 2026 · 1 publisher

  27. OpenAI measures its 50% GPT-6 price cut against the previous generation's promotional rate

    Science · September 22, 2026 · 2 publishers

  28. Bessent announces further talks on a US-China AI incident hotline before Trump meets Xi

    Invest · September 22, 2026 · 1 publisher

  29. Anthropic's worked example turns a 120,000-token conversation into 2.8 million billed input tokens

    Invest · September 22, 2026 · 1 publisher

  30. Xiaomi's MiMo-V2.6-Pro leads the open-weight index at $0.87 per million output tokens

    Product · September 22, 2026 · 1 publisher

  31. Claude Code's per-repository memory store lost a rule set for all projects

    Build · September 21, 2026 · 1 publisher

  32. Attackers are bypassing authentication on Cisco ISE with a crafted API request

    Security · September 21, 2026 · 1 publisher

  33. Antigravity gives away five frontier models on a quota Google keeps trimming

    Product · September 20, 2026 · 1 publisher

  34. Databricks reports 60% higher coding spend on the model benchmarks score as cheaper

    Invest · September 19, 2026 · 1 publisher

  35. Two ordinary defects carried Hacktron from a forum image upload to OpenAI's internal monorepo

    Build · September 19, 2026 · 1 publisher

  36. Word overlap in one listing line decided every skill invocation across 38 Claude Code runs

    Build · September 19, 2026 · 1 publisher

  37. Eight stacked repairs drop ProgramDistill's partial-reconstruction success to 32%

    Build · September 19, 2026 · 1 publisher

  38. Anthropic's follow-up says Mythos 5 acted like a model that knew the internet was real

    Build · September 19, 2026 · 1 publisher

  39. Astra wrote its keepInventory rule only after the Creeper took the chest and the bed

    Build · September 17, 2026 · 2 publishers

  40. Fireworks' own DeepSWE numbers put four coding models inside the noise band

    Product · September 17, 2026 · 1 publisher