GitHub is rolling GPT-6.1 Sol into Copilot, citing early tests in which it used fewer tokens and steps than GPT-6 and GPT-5.6, without figures. Teams will have to measure any saving on their own tasks, at a rate the announcement did not quote.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives55
- Confidence40
One OpenAI Codex prompt spawned 826 child agents and burned about $78,000 in credits, according to the user's own reconstruction. Nearly all the counted tokens trace to an alpha client build, and only OpenAI's servers can turn them into dollars.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence30
Microsoft's new Copilot app bills coding and agents by use, while fewer than 7% of over 450 million commercial 365 seats pay for Copilot. Usage billing lets Microsoft charge its existing Copilot customers more whether or not the unlicensed majority ever signs up.
Reality
- Evidence55
- Adoption15
- Hype gap+15
- Incentives70
- Confidence60
Simon Willison's teardown finds a cloud container whose outbound domain list appears open by default, plus a desktop build that runs programs on the employee's machine. Those are two provisioning reviews with different threat models.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
A dev.to guide to GPT-5.6 pricing shows batch halving both token rates and a cheap-first cascade saving money until 9 in 10 calls escalate. Batch is opt-in and caching fails silently on short prefixes, so the default request often pays list price.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives75
- Confidence50
Tuesday's model releases from Anthropic and OpenAI came with audit numbers that move in both directions at once. For anyone granting an agent tokens or tooling, a version bump means re-running the injection tests.
Reality
- Evidence54
- Adoption30
- Hype gap+18
- Incentives79
- Confidence44
OpenAI's guide prices reused input tokens at a discount of up to 90 percent. It also says cached key-value states sit on individual machines, so an unchanged prefix can still miss the cache when routing overflows.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives75
- Confidence65
Tencent's Hunyuan Speech team put a swappable agent model behind a conversation model that re-decides every second whether to talk, and reports best-in-test timing on Full-Duplex-Bench v3 with a task-accuracy shortfall it describes only as slight.
Reality
- Evidence34
- Adoption12
- Hype gap+18
- Incentives62
- Confidence45
XDA Developers handed Claude Code, Codex and Google Antigravity the same demanding site brief with no follow-ups. All three shipped every section that was asked for. The distance between them showed up in typography, spacing and mobile mockups.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+24
- Incentives40
- Confidence45
California's transparency law makes a developer's own safety policy enforceable against it, and the Midas Project built its case entirely from OpenAI's Frontier Governance Framework and the system cards published after it.
Reality
- Evidence64
- Adoption45
- Hype gap+18
- Incentives70
- Confidence62
Commerce and the White House put three named frontier models from the two leading US labs under access restrictions in June 2026. Enterprise access to those models now runs through a government approval step.
Reality
- Evidence16
- Adoption
- Insufficient
- Hype gap+46
- Incentives62
- Confidence21
Jensen Huang says OpenAI's Astra reached AGI, and the tape answered by putting 15% into CoreWeave and taking 2% out of Nvidia, which says rather more about a $108 billion revenue guide than about machine cognition.
Reality
- Evidence55
- Adoption38
- Hype gap+58
- Incentives80
- Confidence55
The same guide that tells you to run evals on every change now carries a deprecation notice with two dates. For anyone whose deploy gate creates eval runs, the earlier date is the one that bites.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap−28
- Incentives58
- Confidence68
OpenAI's create response reference is explicit that a chained request does not inherit earlier instructions, and because the field is optional, the turn that loses your JSON contract or your safety text still succeeds.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+6
- Incentives32
- Confidence54
SemiAnalysis counts roughly $350m of state money behind five competing Korean foundation-model consortiums, set against a $919bn infrastructure figure. The gap is where the investment case sits, and it points at hardware.
Reality
- Evidence42
- Adoption48
- Hype gap+28
- Incentives55
- Confidence44
Anthropic's revenue keeps climbing, on the FT's sourcing, while the card-billing index Ramp builds from 70,000 companies shows July's Anthropic spend going to cheaper Claude models instead of the flagship.
Publishers:feeds.simonwillison.net
Reality
- Evidence55
- Adoption68
- Hype gap+20
- Incentives62
- Confidence57
Responses swaps runs and threads for input items and conversations. The guide files history pruning, retries and the tool loop under separation of concerns, which means your application owns them now, and prompts move to the dashboard.
Reality
- Evidence74
- Adoption
- Insufficient
- Hype gap+18
- Incentives82
- Confidence70
Fast Company's three-bucket sort of proprietary, open weight and open source works best as a procurement checklist, because the middle bucket hands over the weights and keeps the training corpus out of sight.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+14
- Incentives44
- Confidence55
SemiAnalysis puts 40-50% of 2027's new capacity under OpenAI and Anthropic contracts, with revenue per megawatt now well clear of deployment cost. That makes inference a residual market.
Reality
- Evidence20
- Adoption30
- Hype gap+55
- Incentives65
- Confidence26
Thibault Sottiaux says ChatGPT Work packages Codex for non-engineers at $20 a month. The interview's own two billed themes, discovery and the cost of intelligence, come out unanswered.
Reality
- Evidence28
- Adoption38
- Hype gap+38
- Incentives82
- Confidence46
Earlier coverage
- Same-day GPT-5.6 on Azure kills the parity argument, leaving auth and residency to decide
Build · August 25, 2026 · 1 publisher
- A Government Switched Off Two Frontier Models. Your Board Will Want The Fallback Plan.
Leadership · August 21, 2026 · 1 publisher
- Bedrock turns GPT-5.6 throughput into a routing choice, with residency as the price
Build · August 20, 2026 · 1 publisher
- NVIDIA is documenting Holoscan for the agent, not the engineer
Build · August 20, 2026 · 1 publisher
- GPT-5.6 ships as three models, and that makes model choice a deployment decision
Build · August 18, 2026 · 1 publisher
- Your inference bill is an architecture defect: declare the task before you call the model
Build · August 18, 2026 · 1 publisher
- Developer habit, priced at $965B: what Anthropic's run actually proves
Build · August 15, 2026 · 1 publisher
- Washington's secret AI test is coming for open weights, and release dates go with it
Product · August 14, 2026 · 2 publishers