Skip to content

Topic

Agent Observability

Instrumenting LLM agents with logs, session traces, latency and token metrics, and evals so runs can be inspected and debugged after the fact.

Current stories

product1 publisher

Claude Statuspane makes agent spend visible to one developer at a time

Anji Xu's open-source Claude Statuspane puts context use, five-hour and seven-day rate limits and session cost in a card above the Claude Code prompt. The readout lives on that one developer's screen, so whoever answers for a team's total agent spend still needs records kept somewhere central.

Publishers:devops.com

Reality

Evidence45
Adoption5
Hype gap0
Incentives
Insufficient
Confidence50
build1 publisher

LiteLLM's Lens moves agent-failure analysis into the self-hosted gateway that routes model calls

LiteLLM launched Lens on September 30, a tool that uses AI agents to find recurring failures across agent traces sent through its model gateway. Customers host the analyzer and its databases, and the 200,000-trace volume CTO Ishaan Jaffer cites is a future target Lens has not been measured against.

Publishers:runtimewire.com

Reality

Evidence40
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence40
build1 publisher

Grading one agent session on five dimensions reloads the same trace five times

Warp's case for LLM-as-a-judge scoring is that a coding agent leaves a complete record you can grade after the fact. Each dimension gets a prompt, a rubric and a judge model of its own, and it runs at about 3% of the company's own token bill.

Publishers:warp.dev

Reality

Evidence32
Adoption22
Hype gap+24
Incentives84
Confidence56

Earlier coverage

  1. A Stop hook clocked Claude Code's code-reviewer subagent at 37 seconds per call

    Build · September 10, 2026 · 1 publisher

  2. Anthropic caught six unauthorized agent runs by re-reading 141,006 evaluation logs

    Build · September 2, 2026 · 1 publisher

  3. Seven MCP tool-server bugs billed Databricks $499K a year in retried tokens

    Build · September 1, 2026 · 1 publisher

  4. Bedrock's managed agentic retrieval nests a second loop inside the call your RAG logs count as one

    Build · August 31, 2026 · 1 publisher

  5. An unsupervised agent loop billed $38 before anything in the system said stop

    Build · August 30, 2026 · 1 publisher

  6. A gate that stops firing shows up in PlannerCritic's metrics as safer plans

    Build · August 29, 2026 · 1 publisher

  7. Overflowing Claude Code's skill listing strips the descriptions the model triggers on

    Build · August 28, 2026 · 1 publisher

  8. Natera's voice scheduler: the hard parts were sockets, filler speech and when to ask for ID

    Build · August 26, 2026 · 1 publisher

  9. Agent traces became product data, and the write pattern now picks your storage

    Build · August 26, 2026 · 1 publisher

  10. Four agents, five stages, one manifest row: AWS's migration pipeline is a handoff problem

    Build · August 24, 2026 · 1 publisher

  11. Agents denied a fact do not stop, and read traces cannot tell you they lied

    Build · August 24, 2026 · 1 publisher

  12. AWS's own agent fleet guidance puts the lock-in in state, auth and telemetry, not the framework

    Build · August 24, 2026 · 1 publisher

  13. 157 agent runs, 18 configurations, and the one variable nobody actually tested

    Build · August 23, 2026 · 1 publisher

  14. The weights never moved: what 6,852 Claude Code sessions say about where regressions live

    Build · August 23, 2026 · 1 publisher

  15. ZizkaDB bets agent debugging on edges you declare, not spans you read

    Build · August 23, 2026 · 1 publisher

  16. Green means it did not crash: scheduled agents need an artifact, not an exit code

    Build · August 22, 2026 · 1 publisher

  17. Agents Are Not Microservices With an LLM Attached, and the Retrofit Never Arrives

    Product · August 19, 2026 · 1 publisher

  18. Your agent traces are append-only, which is why they hide the bug

    Build · August 19, 2026 · 1 publisher

  19. An empty array is a claim about your query: verify identifiers before you trust the metric

    Build · August 19, 2026 · 1 publisher

  20. Artificial Analysis moves eval onto your data, and turns model choice into procurement

    Build · August 15, 2026 · 1 publisher

  21. Instrumentation Is the Whole Gap Between an Agent and an Agent You Can Run

    Build · August 15, 2026 · 1 publisher