Anthropic's open SKILL.md format is read by roughly forty agent tools, among them Codex, Cursor, Copilot and Gemini CLI. Each installed skill costs about 100 tokens a session until a task matches its description and the full workflow loads.
Reality
- Evidence45
- Adoption45
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
Platinum River Innovations told theCUBE it moved from time-and-materials to fixed bids and took the delivery risk onto itself. The wider repricing claim made at the Certinia event rests on that one named firm.
Reality
- Evidence30
- Adoption38
- Hype gap+45
- Incentives85
- Confidence45
With documentation and web search removed, GPT-5.6 Luna passed 18% of 336 Dev Proxy tasks and 15% of 413 SPFx tasks, and the passes appear throughout both product histories instead of stopping at one release.
Publishers:devblogs.microsoft.com
Reality
- Evidence62
- Adoption18
- Hype gap+12
- Incentives55
- Confidence55
Dhravya Shah spent three days probing the personal agent from the outside and reports git-tracked Markdown searched by keyword, with a background pass he clocked taking 23 hours and 16 minutes to commit one preference.
Reality
- Evidence42
- Adoption24
- Hype gap+28
- Incentives76
- Confidence58
Justin Johnsen has been a forward-deployed engineer at KPMG for about a year. His account of one engagement puts months of requirements and context work ahead of a build that took roughly a month, then coaching the client's own team.
Reality
- Evidence30
- Adoption25
- Hype gap+15
- Incentives65
- Confidence45
A dev.to piece on harness engineering splits agent degradation into five separable failure modes. The two it actually describes both turn on where a token sits in the assembled prompt, and a bigger window leaves both in place.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+30
- Incentives68
- Confidence55
In a dev.to walkthrough, miruky uses one bad rstrip call to separate five layers of agent engineering. Which layer do you change when a repair passes all three examples and still accepts 1e3s?
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap−5
- Incentives20
- Confidence58
shinpr has taken claude-code-workflows through 133 releases, and the recent ones delete structure the models no longer need. His session reader then found three mandatory steps in his own repository that never ran.
Reality
- Evidence42
- Adoption18
- Hype gap+12
- Incentives58
- Confidence46
Ajay Prakash's InfoQ talk walks a pager alert through logs, a downstream service and a buggy pull request. Each hop in that path depends on debugging instructions another team at LinkedIn wrote down.
Reality
- Evidence44
- Adoption57
- Hype gap+22
- Incentives41
- Confidence56
A field-notes post on production LLM agents argues the context window behaves like a cache under eviction pressure, and that the real work is deciding every turn what earns space and which failed turns to scrub.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives40
- Confidence48
Son Nguyen says his team ran the same prompt and the same use case through different models, found the difference surprisingly small, and rebuilt the preprocessing, context and validation around the model instead. He calls that surrounding system the agent harness.
Reality
- Evidence24
- Adoption15
- Hype gap+32
- Incentives78
- Confidence45
Newer models ship with advice to keep the guidance file under 200 lines. The analyst who wrote those 1,042 lines argues each one logs context the model lacked, and that the real debt is a mechanical rule left sitting in prose.
Reality
- Evidence42
- Adoption22
- Hype gap−12
- Incentives38
- Confidence55
The harness lost its hidden system prompt, 43% of its builtin tool descriptions and its todo list middleware. LangChain's own footnote says reward confidence intervals span zero for every model tested, so the evals settle the token saving more firmly than the quality.
Publishers:langchain.com
Reality
- Evidence58
- Adoption30
- Hype gap+18
- Incentives82
- Confidence46
Polylane moved triage, investigation and code generation into a single run on September 3. Model spend per pull request fell from $111 to about $18, though the agent now files seven times as many of them.
Reality
- Evidence48
- Adoption30
- Hype gap+27
- Incentives62
- Confidence55
Durable state and sandboxed execution now arrive as platform primitives, so a dev.to guide argues the remaining failure mode is undocumented organizational context. It leans on one arXiv paper and one example.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence40
Alberto Souza says he reviewed none of the generated code. The checking he does do happens before any code exists, in context files and successive rounds of questions from the agent that he answers one by one.
Reality
- Evidence28
- Adoption22
- Hype gap+34
- Incentives55
- Confidence52
Writing on dev.to, Ben McCarthy describes rebuilding the guest-messaging agents across his 23 holiday lets around a typed per-request view of state, with world facts held apart from the policy about what the agent may promise.
Reality
- Evidence32
- Adoption20
- Hype gap+15
- Incentives35
- Confidence55
Anton Brilliantov spent eight parts specifying agent handoffs down to a single acceptance command. His ninth names the task shapes where those facts are still unknown, and the reading that has to happen first.
Reality
- Evidence32
- Adoption10
- Hype gap−10
- Incentives25
- Confidence45
A dev.to post hands four models a file whose comment contradicts its code, then asks a clean session to fix the inconsistency. It is a well-built way to expose the failure, and no counts are published yet.
Reality
- Evidence38
- Adoption12
- Hype gap+10
- Incentives22
- Confidence45
Anthropic has cut 80% of Claude Code's system prompt. PostHog's audit of its own agent config file suggests the maintenance job now runs the other way, toward subtraction. A stale line is the cautionary case.
Publishers:newsletter.posthog.com
Reality
- Evidence42
- Adoption38
- Hype gap+18
- Incentives66
- Confidence45
Earlier coverage
- TOON's 49% character saving falls to 33% against JSON that was already minified
Build · August 31, 2026 · 1 publisher
- Testing a skill means running the scenario again on the next model version
Build · August 31, 2026 · 1 publisher
- After five months behind main, classifying 312 conflict hunks helped turn a two-week rebase estimate into 11 hours
Build · August 30, 2026 · 1 publisher
- An unsupervised agent loop billed $38 before anything in the system said stop
Build · August 30, 2026 · 1 publisher
- A memory layer beat CLAUDE.md by 22.2 points, and 54 of 72 test pairs never moved
Build · August 24, 2026 · 1 publisher
- Agents denied a fact do not stop, and read traces cannot tell you they lied
Build · August 24, 2026 · 1 publisher
- Unity's AI problem is not the prompt: the load-bearing context lives in the prefabs
Build · August 24, 2026 · 1 publisher
- Coding agents cost $4,125 a month because 73% of it is context you already sent
Build · August 23, 2026 · 1 publisher
- Anthropic cut 80% of Claude Code's system prompt and the evals did not move
Invest · August 23, 2026 · 1 publisher
- The fix for a confused coding agent is a smaller input, not a bigger window
Build · August 23, 2026 · 1 publisher
- OpenClaw makes the channel the architecture, and the reasoning loop a lodger
Build · August 23, 2026 · 1 publisher
- Agent memory that learns from wins is grading the user, not the context
Build · August 23, 2026 · 1 publisher
- The personal agent is a folder, not a model: four files and less memory than you thought
Product · August 22, 2026 · 1 publisher
- Open Knowledge Format: when the fact already has a name, the chunker is the bug
Build · August 22, 2026 · 1 publisher
- Model choice is becoming a line item, and the differentiator moved up the stack
Product · August 22, 2026 · 1 publisher
- Netflix's plain-text recommender won on 40x fewer labels, and the bill moved rather than vanished
Build · August 22, 2026 · 1 publisher
- Stage Gates Got Cheap Again, And That Is The Whole Argument For "Waterfall 2.0"
Product · August 21, 2026 · 1 publisher
- Three tools, three spellings of the same glob: agent rules do not port
Build · August 20, 2026 · 1 publisher
- Pocock's /wayfinder bets that the bottleneck in overnight agents is planning, not code
Build · August 20, 2026 · 1 publisher
- JFrog measured 847 log lines to find 9, and that ratio is now a budget line
Build · August 20, 2026 · 1 publisher
- Adronite's Codistry makes token count, not context window, the axis of competition
Product · August 19, 2026 · 2 publishers
- Agent memory has a dose-response curve, and the cheapest dose won the biggest gain
Build · August 18, 2026 · 1 publisher
- The reason your agent gets worse after an hour is that nothing ever leaves the context window
Build · August 18, 2026 · 1 publisher
- Your inference bill is an architecture defect: declare the task before you call the model
Build · August 18, 2026 · 1 publisher
- The payload is rebuilt every turn, so stop treating your prompt as a shipped artifact
Build · August 15, 2026 · 1 publisher
- Context rot at 15 iterations: two toolkits that move the spec into Git
Build · August 15, 2026 · 1 publisher
- Agent reliability is a harness problem, not a prompt problem
Build · August 15, 2026 · 1 publisher