Skip to content

model

Claude Opus 4.8

Anthropic model, run with Extra-High Reasoning, used for the study's main agent experiments.

Known aliases

  • claude-opus-4-8
  • claude-opus-4.8
  • Claude Opus 4.8 Extra-High Reasoning
  • Claude Opus 4.8 Fast
  • Claude Opus 4.8 Fast mode
  • MODEL_L3
  • Opus 4.8
  • Opus-4.8

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Moonshot publishes open weights for its trillion-parameter K2.7-Code model

Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence40
build1 publisher

Andon Labs opens the agent platform behind its two money-losing shops

Andon Labs opened Pion, its platform for agent-run companies, as a research preview on September 14, with its own agent-run store and cafe still losing money. For anyone building long-running agents, the shops show what the loop does once it has to pay rent and wages.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+45
Incentives60
Confidence40
build5 publishers

GLM-5.3 keeps GLM-5.2's base model and claims 50% more on coding: plan for shorter eval cycles

Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.

Perspective Coverage

5 publishers
Builder
Builder 58%
Operator
Operator 33%
Investor
Investor 9%

Reality

Evidence40
Adoption30
Hype gap+35
Incentives70
Confidence55
security4 publishers

NASA's AIT-GUI Ground Console Shipped Without Auth: CVSS 9.4, Fixed in 2.5.2

A flaw in NASA's open-source AIT-GUI lets unauthenticated requests reach spacecraft command routes, and Cycode says a malicious web page can deliver them through an operator's browser.

Perspective Coverage

4 publishers
Builder
Builder 43%
Operator
Operator 50%
Investor
Investor 7%

Reality

Evidence70
Adoption
Insufficient
Hype gap+30
Incentives45
Confidence65
product7 publishers

Box measured Claude Opus 5.5 using a third of the tokens Opus 5 needed

Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.

Perspective Coverage

7 publishers
Builder
Builder 41%
Operator
Operator 38%
Investor
Investor 21%

Reality

Evidence45
Adoption30
Hype gap+35
Incentives60
Confidence60
leadership3 publishers

A Friday letter from Commerce turned frontier-model routing into an export-control problem

Commerce told Anthropic on June 12 that any foreign national, anywhere, needs a BIS license to use Fable 5 or Mythos 5. Anthropic disabled both models for every customer to comply, and the letter has not been made public.

Publishers:anthropic.comcsis.orgnatlawreview.com

Perspective Coverage

3 publishers
Builder
Builder 32%
Operator
Operator 41%
Investor
Investor 27%

Reality

Evidence61
Adoption77
Hype gap+9
Incentives74
Confidence66

Earlier coverage

  1. Filling GLM-5.3-Flash's million-token window costs three times its per-task benchmark price

    Build · September 20, 2026 · 2 publishers

  2. Export controls took Anthropic's newest models offline three days after launch

    Leadership · September 16, 2026 · 1 publisher

  3. LangChain drops about 4,000 base input tokens from every default Deep Agents turn

    Build · September 15, 2026 · 1 publisher

  4. A three-agent CrewAI run spent 44.6 of its 118 seconds inside coworker tool calls

    Build · September 14, 2026 · 1 publisher

  5. DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores

    Invest · September 14, 2026 · 1 publisher

  6. Andon Labs' Pion hands one agent the email, phone, banking and cards of a whole business

    Build · September 14, 2026 · 1 publisher

  7. Anthropic's most capable model built working exploits from 16 of 39 published patches

    Leadership · September 11, 2026 · 1 publisher

  8. Stripped GLM-5.3-Flash weights show what Z.ai's MIT license permits

    Build · September 8, 2026 · 1 publisher

  9. GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor

    Leadership · September 5, 2026 · 1 publisher

  10. Three lines of context in the LSP reply cut follow-up file reads from 15.2 to 3.2

    Build · September 3, 2026 · 1 publisher

  11. An attack harness closed 67 points of Booz Allen's own AI threat ranking

    Product · September 3, 2026 · 1 publisher

  12. Pandex hooked a Fortune 500 agent four minutes after claiming a package name from llms.txt

    Build · September 2, 2026 · 1 publisher

  13. Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8

    Build · August 31, 2026 · 1 publisher

  14. Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test

    Build · August 31, 2026 · 2 publishers

  15. Microsoft retired four models from the Foundry router under every deployment left on defaults

    Build · August 31, 2026 · 1 publisher

  16. Ramp's July card data puts Opus 4.8 at 3.5 times Claude Fable's spend share

    Product · August 30, 2026 · 1 publisher

  17. Open weights take 29% of gateway tokens on a twenty-fifth of the dollars

    Invest · August 30, 2026 · 1 publisher

  18. Anthropic's own monitor caught its agents gaming 39 of 1,601 alignment runs

    Invest · August 28, 2026 · 1 publisher

  19. Thomson Reuters spent $40 million to own the layer above the open weights

    Leadership · August 28, 2026 · 1 publisher

  20. Same price, cheaper fast mode: Opus 4.8 argues on unit economics

    Leadership · August 26, 2026 · 1 publisher

  21. Ox Alpha was GLM-5.3-Flash, and the number that decides displacement is 18 billion

    Product · August 26, 2026 · 1 publisher

  22. Anthropic ships a price dial with its new model, and that is now the buying decision

    Leadership · August 26, 2026 · 1 publisher

  23. Google's legal AI bundle lands a day after a $40M model, and the connector list tells you why

    Build · August 26, 2026 · 2 publishers

  24. Grayscale's ZCSH puts Zcash in brokerage accounts, and none of its ZEC is shielded

    Invest · August 25, 2026 · 3 publishers

  25. Claude Code tells you the model, not the culprit: 106 lines of shell to name the Skill

    Build · August 25, 2026 · 1 publisher

  26. Thomson Reuters trades Claude for a Qwen derivative it cannot let customers audit

    Invest · August 25, 2026 · 1 publisher

  27. App factory or agent fleet manager: the fork is whose rate limit stops the work

    Build · August 24, 2026 · 1 publisher

  28. Thomson Reuters priced the middle path at $40M, and still pays Anthropic

    Build · August 24, 2026 · 4 publishers

  29. Four Claude models, four surfaces, one incident: tier fallback is inside the blast radius

    Product · August 24, 2026 · 1 publisher

  30. The AI boss forgot its own handbook, and humans had to hand it back

    Build · August 23, 2026 · 1 publisher

  31. Tier the models; the validation boundary is the thing you are actually buying

    Build · August 22, 2026 · 1 publisher

  32. A 27B-parameter agent beat two frontier models at one task, and the task was chosen carefully

    Product · August 22, 2026 · 1 publisher

  33. Meta's coding agent has two prices: pay 18x more, or let it train on your repository

    Invest · August 22, 2026 · 1 publisher

  34. Physics-only world models cannot predict people, and the fix costs six pipeline stages

    Build · August 22, 2026 · 1 publisher

  35. Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures

    Build · August 21, 2026 · 1 publisher

  36. A Government Switched Off Two Frontier Models. Your Board Will Want The Fallback Plan.

    Leadership · August 21, 2026 · 1 publisher

  37. TrueFoundry open-sources an agent harness and calls managed agents a lock-in play

    Build · August 19, 2026 · 2 publishers

  38. Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design

    Build · August 19, 2026 · 2 publishers

  39. Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.

    Product · August 19, 2026 · 1 publisher

  40. Adronite's Codistry makes token count, not context window, the axis of competition

    Product · August 19, 2026 · 2 publishers