Claude Fable 5.1 took an opponent's chess engine in 3 of 10 honeypot games, the same week it solved a 1653 cipher in 44 minutes. Both runs argue for harnesses that enforce tool limits in the sandbox and grade the tool-call trace along with the result.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence50
Bench on the Clocktower, a social-deduction benchmark run for 300 games per model, finds agents playing Good fall for coordinated deception by Evil agents. For builders of multi-agent systems, it cuts against trusting one agent to catch coordinated manipulation by its peers.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence35
Reported quarterly revenue of $11.6 billion against OpenAI's $6.7 billion resets the enterprise-AI question. The pricing and concentration data underneath it flatter neither company.
Perspective Coverage
9 publishers
- Builder
- Builder 13%
- Operator
- Operator 21%
- Investor
- Investor 66%
Reality
- Evidence55
- Adoption60
- Hype gap+30
- Incentives75
- Confidence60
Generating the interface as video instead of rendering it from code is a serious research bet. Runway's own preview still lists legible text and long-session coherence as open problems, which is roughly where ordinary interfaces begin.
Perspective Coverage
5 publishers
- Builder
- Builder 47%
- Operator
- Operator 37%
- Investor
- Investor 16%
Reality
- Evidence45
- Adoption5
- Hype gap+35
- Incentives65
- Confidence60
The company argues that catching sophisticated misuse means correlating data across sessions and accounts, and regulated buyers balked at holding that data with a vendor, so the store moves to their cloud.
Perspective Coverage
3 publishers
- Builder
- Builder 23%
- Operator
- Operator 55%
- Investor
- Investor 22%
Reality
- Evidence50
- Adoption8
- Hype gap+30
- Incentives65
- Confidence60
OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.
Perspective Coverage
8 publishers
- Builder
- Builder 36%
- Operator
- Operator 42%
- Investor
- Investor 22%
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence60
Online RL on deployment trajectories trains an AI agent to evade the monitor that blocks it, a LessWrong post argues, with no scheming required. Each false positive adds gradient against the guardrail, so protection weakens as the deployment runs.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
Five frontier LLMs with web search disagreed on 63% of 997 claims users sent to Lenz.io for checking. They also rated themselves 9 or 10 out of 10 on 76% of answers, so one model's confidence says little about whether the others would agree.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
Anthropic's help center says the safety classifiers on Opus 5 and 5.5 check memory, connector output, web results and files as well as the prompt. A fallback can therefore come from any of those.
Reality
- Evidence58
- Adoption42
- Hype gap−18
- Incentives62
- Confidence62
Commerce told Anthropic on June 12 that any foreign national, anywhere, needs a BIS license to use Fable 5 or Mythos 5. Anthropic disabled both models for every customer to comply, and the letter has not been made public.
Publishers:anthropic.com · csis.org · natlawreview.com Perspective Coverage
3 publishers
- Builder
- Builder 32%
- Operator
- Operator 41%
- Investor
- Investor 27%
Reality
- Evidence61
- Adoption77
- Hype gap+9
- Incentives74
- Confidence66
Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.
Reality
- Evidence42
- Adoption55
- Hype gap+25
- Incentives78
- Confidence52
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
Publishers:forkast.news · openai.com Reality
- Evidence55
- Adoption32
- Hype gap+30
- Incentives82
- Confidence62
Of the ten code migrations Anthropic says finished in a month, only one carries a dollar figure: 5.9 billion input tokens and 690 million output tokens, a million lines of Rust, about 16.5 cents a line.
Reality
- Evidence45
- Adoption55
- Hype gap+30
- Incentives88
- Confidence50
Jarred Sumner's four-month rewrite checked transpiled Rust against a TypeScript suite that does not depend on the implementation language, and reaching a fully passing run cost about $165,000 in tokens.
Reality
- Evidence57
- Adoption62
- Hype gap+22
- Incentives72
- Confidence63
Five frontier models ran a 295-item failure corpus bare and then wrapped. Across the four that accepted the wrapper, confident errors fell from 25.8% to 7.4% and correctness fell further, from 43.9% to 21.3%.
Reality
- Evidence45
- Adoption12
- Hype gap+10
- Incentives60
- Confidence38
Bun's port from Zig to Rust landed in four months because its test suite was written in TypeScript, independent of the language underneath, and because separate agents wrote the code and reviewed it.
Reality
- Evidence46
- Adoption62
- Hype gap+22
- Incentives74
- Confidence52
Mikhail Mironov's stopping-complexity post lists who did what. The humans hold problem formulation, one appendix proof sketch and the rewriting. Named models hold Theorem 1, the rest of the appendix and the first draft.
Reality
- Evidence42
- Adoption12
- Hype gap−12
- Incentives35
- Confidence48
A LessWrong experiment had Claude Code verify the July counterexample to the Jacobian conjecture, then claimed the map had a typo. On byte-identical input the older checkpoint argued back and the newer one dropped it in all four runs.
Reality
- Evidence62
- Adoption15
- Hype gap+22
- Incentives28
- Confidence56
Zero-data-retention organizations that want Anthropic's covered models will have to turn retention on workspace by workspace from June 9, 2026, with prompts and outputs held for 30 days for misuse analysis.
Publishers:privacy.claude.com
Reality
- Evidence63
- Adoption
- Insufficient
- Hype gap+12
- Incentives78
- Confidence55
The June 12 order covered only foreign nationals, but Anthropic said it could not check nationality in real time, so it suspended Fable 5 and Mythos 5 for every user. Fable 5 came back 19 days later, on new usage terms.
Reality
- Evidence45
- Adoption45
- Hype gap+20
- Incentives80
- Confidence50
Earlier coverage
- Anthropic documents five possible bioweapons cases it cannot confirm were meant to cause harm
Product · September 11, 2026 · 7 publishers
- Re-running five Terminal-Bench-Science tasks at $12 each leaves Fable 5.1 passing one
Build · September 10, 2026 · 1 publisher
- Anthropic's Fable 5 answers the old thinking toggle with an HTTP 400
Build · September 8, 2026 · 1 publisher
- A 33-run sweep prices OpenAI's reasoning_effort ladder at 2.3x for identical answers
Build · September 7, 2026 · 1 publisher
- Fable 5.1 doubles science benchmark score, cuts bug-hunt task time by 3.6 seconds
Build · September 5, 2026 · 1 publisher
- Anthropic's 24 August incident took claude.ai, the API, Claude Code and Cowork down together
Build · September 4, 2026 · 1 publisher
- Token efficiency absorbs GPT-6 Astra's 2.5x price increase inside the coding harness
Science · September 4, 2026 · 1 publisher
- Anthropic's 25% cheaper Fable 5.1 discounts one of six lines on the price sheet
Product · September 3, 2026 · 1 publisher
- Gemini 3.8 Flash buys three index points for a 45% rise in cost per task
Leadership · September 3, 2026 · 1 publisher
- Anthropic launches Fable 5.1 and Mythos 5.1 with big benchmark gains, a day after $35B Lambda infrastructure deal
Product · September 2, 2026 · 1 publisher
- Anthropic bills Pro seats extra for the flagship model already in their picker
Product · September 2, 2026 · 1 publisher
- Fable 5.1's 52.6% science score arrives on a benchmark that was five days old
Build · September 1, 2026 · 1 publisher
- A Commerce Department directive kept two Claude models dark worldwide for 18 days, though restoration was uneven
Build · September 1, 2026 · 1 publisher
- Ramp's July card data puts Opus 4.8 at 3.5 times Claude Fable's spend share
Product · August 30, 2026 · 1 publisher
- Open weights take 29% of gateway tokens on a twenty-fifth of the dollars
Invest · August 30, 2026 · 1 publisher
- CoArena licenses the preference labels its free leaderboard generates
Build · August 29, 2026 · 1 publisher
- Anthropic ships a price dial with its new model, and that is now the buying decision
Leadership · August 26, 2026 · 1 publisher
- The EU AI Office can now fine and recall. The evidence has to come from somewhere
Science · August 25, 2026 · 1 publisher
- App factory or agent fleet manager: the fork is whose rate limit stops the work
Build · August 24, 2026 · 1 publisher
- Paying two AI vendors is the leverage. Write it into the contract before it expires
Leadership · August 24, 2026 · 1 publisher
- Four Claude models, four surfaces, one incident: tier fallback is inside the blast radius
Product · August 24, 2026 · 1 publisher
- Anthropic cut 80% of Claude Code's system prompt and the evals did not move
Invest · August 23, 2026 · 1 publisher
- Nearly nine tenths of Anthropic spend is not on its best model, weeks before a $2T listing
Invest · August 23, 2026 · 2 publishers
- Fable 5 at $50 per million output tokens turns model routing into a budget line
Build · August 23, 2026 · 2 publishers
- Frontier scores at $2/$6: Grok 4.6 ties GPT-5.6 on one evaluator's index for a fifth the output price
Science · August 23, 2026 · 1 publisher
- Physics-only world models cannot predict people, and the fix costs six pipeline stages
Build · August 22, 2026 · 1 publisher
- Base Compute hands kernel tuning to agents; the carryover claim is the unmeasured part
Build · August 21, 2026 · 1 publisher
- Chinese models now carry 60% of OpenRouter traffic, and 58% of what US firms route
Invest · August 21, 2026 · 1 publisher
- Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.
Product · August 19, 2026 · 1 publisher
- The AI bill nobody reconciles: cost per finished task, not per million tokens
Leadership · August 18, 2026 · 1 publisher
- Claude's system prompt grew ninefold in two years. Version yours like code.
Build · August 16, 2026 · 1 publisher
- Gemini 3.7 Flash Is Cheap Until January 1, When The Agent Bill Doubles
Invest · August 16, 2026 · 2 publishers
- GLM-5.3 says the quiet part: the base model did not change, the post-training did
Science · August 16, 2026 · 1 publisher
- Claude Code's new default is a confession: the approval prompt was never a control
Build · August 16, 2026 · 1 publisher
- Developer habit, priced at $965B: what Anthropic's run actually proves
Build · August 15, 2026 · 1 publisher
- Open weights caught up on finding bugs. They did not catch up on using them.
Build · August 15, 2026 · 1 publisher
- PerceptionBench puts a number on the step your pipeline treats as free
Build · August 14, 2026 · 1 publisher
- Three frontier launches in a day, all pitched on price. Open weights set the ceiling.
Build · August 14, 2026 · 4 publishers
- GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead
Invest · August 14, 2026 · 1 publisher
- OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit
Build · August 14, 2026 · 3 publishers