Skip to content

Topic

Agent Evaluation as Cross-Layer Measurement

An evaluation approach that checks each AI agent control layer—prompt, context, harness, loop, and graph—rather than judging only final output quality.

Current stories

Earlier coverage

  1. LangChain drops about 4,000 base input tokens from every default Deep Agents turn

    Build · September 15, 2026 · 1 publisher

  2. A production agent's LLM judge filed 113 of 114 state defects under brand voice

    Build · September 13, 2026 · 1 publisher

  3. Procedural Graphs keep an agent's flowchart edit only after it passes a validation set

    Build · September 10, 2026 · 1 publisher

  4. Werewolf agents read the harness's own turn order as evidence of guilt

    Build · September 10, 2026 · 1 publisher

  5. AWS's turn-level metric separates the one broken turn from the three that inherited it

    Build · September 10, 2026 · 1 publisher

  6. Two canaries with known outcomes tell you whether an agent eval is scoring the model or the harness

    Build · September 9, 2026 · 1 publisher

  7. Anthropic's multi-agent writeup puts the engineering weight on coordination and evaluation

    Leadership · August 31, 2026 · 1 publisher

  8. AWS puts agent evaluation on OpenTelemetry, and on exactly three span roles

    Build · August 26, 2026 · 2 publishers

  9. Anthropic cut 80% of Claude Code's system prompt and the evals did not move

    Invest · August 23, 2026 · 1 publisher

  10. Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning

    Build · August 23, 2026 · 1 publisher

  11. The missing field is the product: why agent listings need declared, derived and unknown

    Build · August 23, 2026 · 1 publisher

  12. A SKILL.md layer quietly rerouted an agent off the MCP tools it was given

    Build · August 22, 2026 · 1 publisher

  13. Your agent didn't misunderstand the prompt. It ran the wrong branch.

    Build · August 16, 2026 · 1 publisher