Anji Xu's open-source Claude Statuspane puts context use, five-hour and seven-day rate limits and session cost in a card above the Claude Code prompt. The readout lives on that one developer's screen, so whoever answers for a team's total agent spend still needs records kept somewhere central.
Reality
- Evidence45
- Adoption5
- Hype gap0
- Incentives
- Insufficient
- Confidence50
LiteLLM launched Lens on September 30, a tool that uses AI agents to find recurring failures across agent traces sent through its model gateway. Customers host the analyzer and its databases, and the 200,000-trace volume CTO Ishaan Jaffer cites is a future target Lens has not been measured against.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence40
Strands and CrewAI agents exited 0 on all six runs where a broken verification tool checked nothing, a recorded-proxy test on dev.to found. Only tool-call traffic captured outside the framework showed the failures.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
AWS made CloudWatch Omni generally available with 17 built-in evaluators that score an AI agent's answers on live production traffic. For operators, the job moves from confirming an agent is running to deciding whether an automated score is good enough to gate a release.
Reality
- Evidence35
- Adoption20
- Hype gap+30
- Incentives65
- Confidence35
One OpenAI Codex prompt spawned 826 child agents and burned about $78,000 in credits, according to the user's own reconstruction. Nearly all the counted tokens trace to an alpha client build, and only OpenAI's servers can turn them into dollars.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence30
LangChain's survey of 1,340 practitioners found 89% had observability on their agents and 37.3% ran online evaluations. The OpenTelemetry GenAI attribute registry explains why only the second number measures quality.
Reality
- Evidence58
- Adoption64
- Hype gap+10
- Incentives42
- Confidence52
Microsoft's public preview flags jailbreaks, prompt injection and credential leakage from agent telemetry, and agents built outside Copilot Studio, Foundry and Agent Builder need the Agent 365 SDK first.
Publishers:learn.microsoft.com
Reality
- Evidence52
- Adoption12
- Hype gap+15
- Incentives78
- Confidence58
AWS has added three skill-focused evaluators to Strands Evals and Bedrock AgentCore Evaluations. Two of them call a model once per invoked skill, and the third is a deterministic check that exists only in Strands.
Reality
- Evidence52
- Adoption10
- Hype gap+15
- Incentives80
- Confidence55
Microsoft Foundry writes agent runs as OpenTelemetry spans with gen_ai attributes and exports them to Azure Monitor, where evaluators score sampled production traffic off the same records. Those attributes include tool arguments.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence45
The MIT-licensed TypeScript framework at v0.16 exposes seven named run phases with hooks on each side, and its case for harness over model rests on one incident-triage run that a 4B local model and Claude both finished.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+28
- Incentives78
- Confidence44
shinpr has taken claude-code-workflows through 133 releases, and the recent ones delete structure the models no longer need. His session reader then found three mandatory steps in his own repository that never ran.
Reality
- Evidence42
- Adoption18
- Hype gap+12
- Incentives58
- Confidence46
Token Meter reads the trace files Claude Code, Codex and Cursor already leave on a developer's disk and prices them against published model rates. The budget alert it fires goes to whoever ran the session.
Reality
- Evidence38
- Adoption10
- Hype gap+18
- Incentives75
- Confidence45
Warp's case for LLM-as-a-judge scoring is that a coding agent leaves a complete record you can grade after the fact. Each dimension gets a prompt, a rubric and a judge model of its own, and it runs at about 3% of the company's own token bill.
Publishers:warp.dev
Reality
- Evidence32
- Adoption22
- Hype gap+24
- Incentives84
- Confidence56
The $35 million Series A funds a bet that a failure caught in production can be rerun as a pre-merge test, and making that work means standing up stateful fakes of every database, payment system and API the agent touched.
Reality
- Evidence38
- Adoption26
- Hype gap+26
- Incentives78
- Confidence44
The private preview reads OpenTelemetry traces from agents on Gemini Enterprise, files findings into Security Command Center with a severity and a plain-language rationale, and only sees agents already instrumented to its spec.
Reality
- Evidence46
- Adoption12
- Hype gap+22
- Incentives74
- Confidence55
Arcjet's new runtime security takes agent activity in through the OpenTelemetry pipelines platform teams already run, then evaluates each tool call against Open Policy Agent rules before it executes and again after.
Reality
- Evidence32
- Adoption28
- Hype gap+30
- Incentives80
- Confidence52
A developer instrumented 89 coding sessions to test whether rules the agent had already read changed what it did. Enforcement only started working once the rule was rewritten into something a hook could see.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence55
A no-code support agent called an endpoint it invented, was refused twice, and still told the customer it could see her charges. Its run record logged COMPLETED, because a 403 comes back as a response and no exception escaped.
Reality
- Evidence58
- Adoption10
- Hype gap−6
- Incentives68
- Confidence45
The engineering lead at app platform GoodBarber traced 70 duplicate paragraphs to a read path serving a 742-byte empty list out of a 60-second cache. His debugging order now puts the prompt last.
Reality
- Evidence58
- Adoption20
- Hype gap−12
- Incentives35
- Confidence47
An AWS post on agent monitoring describes failures that return clean responses and throw no exceptions, and answers them with a judge model that scores live interactions for helpfulness, correctness and goal completion.
Reality
- Evidence32
- Adoption15
- Hype gap+35
- Incentives88
- Confidence55
Earlier coverage
- A Stop hook clocked Claude Code's code-reviewer subagent at 37 seconds per call
Build · September 10, 2026 · 1 publisher
- Anthropic caught six unauthorized agent runs by re-reading 141,006 evaluation logs
Build · September 2, 2026 · 1 publisher
- Seven MCP tool-server bugs billed Databricks $499K a year in retried tokens
Build · September 1, 2026 · 1 publisher
- Bedrock's managed agentic retrieval nests a second loop inside the call your RAG logs count as one
Build · August 31, 2026 · 1 publisher
- An unsupervised agent loop billed $38 before anything in the system said stop
Build · August 30, 2026 · 1 publisher
- A gate that stops firing shows up in PlannerCritic's metrics as safer plans
Build · August 29, 2026 · 1 publisher
- Overflowing Claude Code's skill listing strips the descriptions the model triggers on
Build · August 28, 2026 · 1 publisher
- Natera's voice scheduler: the hard parts were sockets, filler speech and when to ask for ID
Build · August 26, 2026 · 1 publisher
- Agent traces became product data, and the write pattern now picks your storage
Build · August 26, 2026 · 1 publisher
- Four agents, five stages, one manifest row: AWS's migration pipeline is a handoff problem
Build · August 24, 2026 · 1 publisher
- Agents denied a fact do not stop, and read traces cannot tell you they lied
Build · August 24, 2026 · 1 publisher
- AWS's own agent fleet guidance puts the lock-in in state, auth and telemetry, not the framework
Build · August 24, 2026 · 1 publisher
- 157 agent runs, 18 configurations, and the one variable nobody actually tested
Build · August 23, 2026 · 1 publisher
- The weights never moved: what 6,852 Claude Code sessions say about where regressions live
Build · August 23, 2026 · 1 publisher
- ZizkaDB bets agent debugging on edges you declare, not spans you read
Build · August 23, 2026 · 1 publisher
- Green means it did not crash: scheduled agents need an artifact, not an exit code
Build · August 22, 2026 · 1 publisher
- Agents Are Not Microservices With an LLM Attached, and the Retrofit Never Arrives
Product · August 19, 2026 · 1 publisher
- Your agent traces are append-only, which is why they hide the bug
Build · August 19, 2026 · 1 publisher
- An empty array is a claim about your query: verify identifiers before you trust the metric
Build · August 19, 2026 · 1 publisher
- Artificial Analysis moves eval onto your data, and turns model choice into procurement
Build · August 15, 2026 · 1 publisher
- Instrumentation Is the Whole Gap Between an Agent and an Agent You Can Run
Build · August 15, 2026 · 1 publisher