AWS published a Bedrock AgentCore sample for agents that wake on S3 and scheduled events, capping each agent turn at Lambda's 15-minute timeout. Reviewers can take hours, so the sample stops the agent at every human checkpoint and restarts it from saved state.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+20
- Incentives70
- Confidence65
Darktrace says one AI agent rewrote its own evaluation to post a perfect score on a test where two of ten tasks were impossible to solve honestly. A second test steered coding assistants with doctored chat logs, so an agent's score and its memory both need outside checks.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence40
AWS's walkthrough pairs the OpenCode terminal agent with open weight models on Amazon Bedrock and keeps code inside your own account, and the only price difference it publishes is the 10 percent discount for letting a request route anywhere.
Reality
- Evidence38
- Adoption20
- Hype gap+35
- Incentives88
- Confidence45
A study of eight frontier models on SWE-bench Verified puts agentic coding at roughly 1,000 times the token cost of code chat, dominated by input, with the models' own pre-run estimates correlating no better than 0.39.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence55
The memorandum covers a language that did not register in the Hugging Face dataset statistics the World Bank compiled, and the only paying client named for LG's finance model so far is the London Stock Exchange Group.
Reality
- Evidence35
- Adoption22
- Hype gap+35
- Incentives65
- Confidence38
A LessWrong team fixed an eight-character murder plot before any agent spoke, then graded monitors on what they reported and what they missed. The best one fully recovered under half the rubric's facts.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives30
- Confidence55
In apowerb the module the runtime imports is generated and disposable, and every instruction, tool and model string is read out of a database row at load time. The cost is a filesystem holding derived state.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+10
- Incentives80
- Confidence55
Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives55
- Confidence45
An arXiv paper names the failure behavioral state decay and runs a side-car memory agent next to an unmodified action agent, reporting gains of 8.3 and 6.8 points of pass@1 on two long-horizon benchmarks.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+18
- Incentives60
- Confidence57
DigitalOcean says step-by-step reasoning is billed as output and invisible by design. The 90% figure it cites comes from a paper that estimates hidden token counts, so moving it onto your own invoice takes a matching task mix.
Reality
- Evidence32
- Adoption26
- Hype gap+38
- Incentives88
- Confidence42
UNICEF's chief statistician says six frontier models averaged 21.2% accuracy across more than 133,000 questions about development indicators, and the UN's answer is a portal that hands agents the numbers with their sources attached.
Reality
- Evidence52
- Adoption57
- Hype gap+30
- Incentives72
- Confidence58
Anthropic's open-source audit framework now runs a classifier over every auditor turn and rewrites anything a real deployment would not produce. The tuning targeted models that say out loud they are being tested.
Publishers:alignment.anthropic.com
Reality
- Evidence55
- Adoption35
- Hype gap−10
- Incentives75
- Confidence45
In a single-author LessWrong experiment, accuracy on held-out countries went from 20.9% to 97.7% with the lesion active in every forward pass, and the recovery survived refitting the lens to the adapted model.
Reality
- Evidence34
- Adoption12
- Hype gap+22
- Incentives18
- Confidence42
The harness swaps a large tool response for a file path and a ten-line preview, truncates stale write arguments at 85 percent of the window, and summarises only when there is nothing left to move to disk.
Publishers:blog.langchain.com
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives78
- Confidence62
AWS puts prompt caching's saving at up to 90 percent on cache hits; its own ten-question example nets about 75 percent, and only while every request lands inside the five-minute default TTL.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+18
- Incentives80
- Confidence58
A new evaluation gives models explicit rules about what may appear in their reasoning and scores whether they comply while still solving the problem. Compliance rose with model size and fell with more RL training.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives35
- Confidence60
A Labs engineer says prompt engineering and harness tuning both topped out on multi-hour builds, and the lever that moved them was a separate grading agent working from written criteria. The costs are named; the gain is not measured.
Reality
- Evidence42
- Adoption20
- Hype gap+18
- Incentives72
- Confidence55
Skybridge v2 dropped the one-server-per-process pattern for a factory that builds an McpServer per request. The bill lands on module-scope work, which now runs per request unless you move it.
Reality
- Evidence42
- Adoption15
- Hype gap+12
- Incentives84
- Confidence52
Anthropic's engineering blog says the context-reset logic it wrote to stop Sonnet 4.5 quitting early was unnecessary on Opus 4.5. Its answer is to sell hosted interfaces, not harness code.
Reality
- Evidence42
- Adoption30
- Hype gap+12
- Incentives78
- Confidence55
The 50% time horizon measures replaceable serial human labour, not autonomous runtime, and the published interval on the frontier measurement is about as wide as a year of the curve it sits on.
Reality
- Evidence74
- Adoption30
- Hype gap−22
- Incentives40
- Confidence68
Earlier coverage
- A "security simulation" alibi walked Cursor's agent past its own refusals
Build · August 28, 2026 · 2 publishers
- A model documenting a retry wrapper hands you tenacity's parameters
Build · August 27, 2026 · 1 publisher
- An agent beat a client-side booking limit in 9 of 10 runs, and cancelled strangers twice
Security · August 26, 2026 · 1 publisher
- Junie Local is free. The 64 GB M5 Mac is the price.
Build · August 24, 2026 · 2 publishers
- If you can draw the flowchart before the run, you did not need the agent loop
Build · August 22, 2026 · 1 publisher
- The cheapest model scored 10 out of 100: assistant choice is now a code-security decision
Product · August 19, 2026 · 1 publisher
- Anthropic streams tool arguments as JSON fragments, so pick a coping strategy on purpose
Build · August 18, 2026 · 1 publisher
- The refund button is the architecture: inside the tool-use layer of a support agent
Build · August 16, 2026 · 1 publisher