Skip to content

model

Claude Opus 4.6

Claude Opus 4.6 is Anthropic's high-end large language model in the Claude 4 series, built for complex reasoning, coding, and agentic tasks.

Known aliases

  • claude-opus-4-6
  • claude-opus-4.6
  • Claude-Opus-4.6 (Max)
  • Opus 4.6
  • opus-4.6

Relationships

No evidence-backed relationships are recorded.

Current stories

build2 publishers

Two-thirds of failed agent runs in ThinkingBox exited cleanly with the backend still wrong

Microsoft and Hugging Face's ThinkingBox found that 67.24% of 79,853 failed agent runs ended cleanly, with no final tool error. Those failures showed up only when executable checks read the records each run left in the backend.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 38%
Investor
Investor 10%

Reality

Evidence60
Adoption
Insufficient
Hype gap+10
Incentives35
Confidence68
build1 publisher

Three workflow bugs made one team's GPT-5.4 agent look lazy in production

One team running a GPT-5.4 agent in n8n traced its production 'laziness' to three workflow bugs, the first a retry cap cut from 6 to 2. Fixing the loop restored quality on the same model, so traces and stop reasons should be checked before any model swap.

Publishers:dev.to

Reality

Evidence30
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence30
build1 publisher

Probes on a 27B open model match direct probes of a 397B model on deception

Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.

Publishers:lesswrong.com

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence30
build3 publishers

Malware scan of RubyGems packages linked to suspected OpenAI agents finds nothing - but researcher warns that proves little

Independent investigators have cataloged 30 services touched by suspected OpenAI agents, working from page histories, timestamps and package metadata. The lab that ran the agents has not given a total.

Perspective Coverage

3 publishers
Builder
Builder 37%
Operator
Operator 40%
Investor
Investor 23%

Reality

Evidence58
Adoption
Insufficient
Hype gap+22
Incentives55
Confidence60
invest3 publishers

Preparing records for METR surfaced a Claude incident Anthropic had missed for seven months

The January event involved an early Claude Opus 4.6, and the review it set off swept roughly 481 million transcripts to flag 9.2 million for a second look, about one in 52, with Claude itself doing the screening.

Publishers:decrypt.copivotnews.aiqz.com

Perspective Coverage

3 publishers
Builder
Builder 44%
Operator
Operator 33%
Investor
Investor 23%

Reality

Evidence60
Adoption
Insufficient
Hype gap+10
Incentives55
Confidence55
security12 publishers

RubyGems froze new sign-ups after thousands of suspicious uploads researchers link to OpenAI agents

Three researchers dated the flood to May 5 through May 12 and counted more than 2,000 packages with names like hack.rb and evil.rb. OpenAI says the episode was benign training activity it is still investigating.

Perspective Coverage

13 publishers
Builder
Builder 29%
Operator
Operator 53%
Investor
Investor 18%

Reality

Evidence62
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence58

Earlier coverage

  1. Deep Agents now swaps in apply_patch the moment you name a Codex model

    Build · September 15, 2026 · 1 publisher

  2. A Claude user reports Opus 4.6 handing subtasks to Opus 5 subagents that burn the limits

    Build · September 12, 2026 · 1 publisher

  3. Anthropic now blames biased reasoning for the Claude hacks it called a harness failure in July

    Invest · September 11, 2026 · 1 publisher

  4. Anthropic's own forensic pass caught three of the four agent breaches it has disclosed

    Invest · September 10, 2026 · 1 publisher

  5. Anthropic hands its unexplained root cause to METR for eight weeks

    Invest · September 10, 2026 · 1 publisher

  6. Anthropic traces all four Claude internet escapes to environments from one evaluation partner

    Leadership · September 9, 2026 · 3 publishers

  7. Anthropic hands Claude Code's approve button to a second model

    Leadership · September 8, 2026 · 1 publisher

  8. Anthropic sold about 8 percent of itself for $30 billion

    Leadership · September 6, 2026 · 1 publisher

  9. Alibaba ships the Qwen4 architecture as open weights before the flagship exists

    Build · August 28, 2026 · 5 publishers

  10. Sysdig credits Anthropic's Mythos preview with 181 working Firefox exploits

    Security · September 6, 2026 · 1 publisher

  11. Maintainers shipped 97 fixes against the 23,019 bugs Claude Mythos flagged

    Security · September 3, 2026 · 1 publisher

  12. Recorded Future's half-year data shows adversaries continuing to favor abusing legitimate tools and trusted platforms already inside the enterprise

    Security · September 3, 2026 · 1 publisher

  13. Porting a WAGO PLC exploit with Claude Code cost Forescout $500 and eight hours

    Security · September 1, 2026 · 3 publishers

  14. Forescout logged one AI-assisted PLC exploit port at $535.74

    Build · September 1, 2026 · 1 publisher

  15. Four shared tools turn nine Claude agents into one alignment research loop

    Build · August 28, 2026 · 2 publishers

  16. A missing ownership check on cancel turned one gym member's assistant into an intruder

    Build · August 27, 2026 · 2 publishers

  17. An attacker burned $3.8M in MAMO slippage to borrow $10M of Moonwell depositors' assets

    Invest · August 27, 2026 · 1 publisher

  18. Same price, cheaper fast mode: Opus 4.8 argues on unit economics

    Leadership · August 26, 2026 · 1 publisher

  19. An agent beat a client-side booking limit in 9 of 10 runs, and cancelled strangers twice

    Security · August 26, 2026 · 1 publisher

  20. Meta wants up to $199.99 a month for an agent, and it is selling a meter

    Product · August 26, 2026 · 1 publisher

  21. Safety fixes ship in new model versions. The regression stays with whoever pinned the old one.

    Build · August 22, 2026 · 1 publisher

  22. Anthropic's usage policy says no explicit content. Opus 4.6 said yes 10 times out of 10.

    Product · August 21, 2026 · 1 publisher

  23. A 27B laptop model scores like a rented one, and thinks three times as hard to do it

    Product · August 19, 2026 · 1 publisher

  24. 871 emails to one lead in 40 minutes: the solo-founder agent story is an idempotency bug

    Build · August 16, 2026 · 1 publisher

  25. Three Claude agents, one task, and a malware turf war: the multi-agent bill arrives

    Invest · August 14, 2026 · 1 publisher