BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Microsoft DevRel's AX playbook traces coding-agent mistakes back to docs, skills and MCP tools
Microsoft DevRel published the AX Practitioner Playbook, a method for finding why coding agents misuse a platform, drawn from hundreds of agent sessions. Its case is that owners of docs, MCP servers and CLIs can fix agent behavior now, if their evals can prove which change helped.
The Engineer · Build desk

What happened
- Microsoft DevRel has measured coding agents on Azure, Cosmos DB, SharePoint Framework and Microsoft 365 Copilot extensions since fall 2025, using prompts like those real developers write.
- The post names docs, MCP tools, skills, plugins, instructions, CLIs and APIs as the agent sources a platform team can actually change.
- The playbook walks through an SPFx project upgrade end to end, from the test scenario through the fixes that shipped.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure Teams already scoring agents with model judges may be reporting passes for code that never compiled, unless a gate actually executes the output.
- decision Before rewriting docs, an SDK or MCP server owner has to establish whether the agent loaded and called their extension, since a content fix cannot reach a skill the agent never opened.
- cost The expensive input is domain experts' time writing, calibrating and versioning criteria, work the playbook says should not be handed to a model.
"An evaluation can lie to you convincingly," Microsoft's post says, and its examples are specific [6]. It reports perfect scores for code that never compiled, and a "used Platform X" check that passed whether or not the agent used Platform X [6]. A check that passes on every input measures nothing. The playbook's evaluation model runs from criteria that judge meaning to gates that prove the code runs [7]. The run gate is the cheapest layer to build, and it would have caught the first of those two failures.
Microsoft calls writing criteria the hardest step [17]. The playbook asks for criteria a judge rules on the same way every time, calibrated before anyone trusts them and versioned when the product changes [8]. It warns against letting a model write them [8]. I think the ordering is correct. A judge that flips its verdict on a rerun cannot tell you whether last week's docs edit helped, and the playbook wants every proposed change tested as a hypothesis before it ships [10].
"A readout tells you what failed while the trajectory tells you why," Microsoft wrote [16]. The split I would adopt first sorts failures into three cases, since each calls for its own fix: the agent never loaded the extension, loaded it but never called it, or called it and applied it incorrectly [9]. In my view the first two are discovery problems. Rewriting the content of a skill the agent never opened changes nothing. The playbook also catalogs nine failure patterns that Microsoft says recur across every technology it evaluated, each pointing to the surface to inspect first [18].
The post's account of a bad generation is blunt: the wrong SDK version, a deprecated authentication pattern, and an agent that "did exactly what its training data told it to do" [15]. It adds that a knowledge cutoff "tells you little about what a model knows about your product" [3]. If the cutoff is no guide, the prior has to be measured product by product. Every lever the post names [4] changes behavior only if the agent reaches it during the session. So the loaded-or-called question comes before any content fix.
The outcome evidence is thinner. Microsoft credits the evaluations with dozens of shipped fixes, including 46 improvements to the Azure Cosmos DB Agent Kit [11]. Forty-six is a count of changes. The post does not report how agent pass rates moved after them. The scenarios ran on Microsoft's own products [5]. For the results to carry to another platform, its developers' prompts have to resemble the scenarios its team writes, and its agents need a route to its extensions at all. The method does not require a particular evaluation system, and the playbook lists the capabilities to check for in whichever one a team uses [13]. Most of the adoption cost is therefore expert time spent on criteria.
Microsoft also released the method as an AX Practitioner skill to install in a coding agent and question during an evaluation [14]. Like any skill, it has to be loaded before it can help.
What to watch
- Whether Microsoft publishes before-and-after agent pass rates for the Cosmos DB Agent Kit or SPFx scenarios, which would turn a count of fixes into an outcome.
- Whether the nine failure patterns hold when teams outside Microsoft run the method on non-Microsoft SDKs and MCP servers.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption15
- Hype gap+10
- Incentives55
- Confidence45
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Microsoft DevRel published the AX Practitioner Playbook, described as its method for evaluating what AI coding agents do with a technology, diagnosing why they get it wrong, and fixing it at the source; it is downloadable at aka.ms/ax-playbook.
- [2]
The playbook puts in one place what Microsoft DevRel learned from hundreds of agent sessions.
- [3]
"Waiting for models to get better isn't a strategy: a knowledge cutoff tells you little about what a model knows about your product."
- [4]
The sources agents rely on that teams can change are docs, MCP tools, skills, plugins, instructions, CLIs, and APIs.
- [5]
Since fall 2025, Microsoft DevRel has measured how agents perform with Azure, Cosmos DB, SharePoint Framework (SPFx), and Microsoft 365 Copilot extensions, using the same kinds of prompts real developers use.
- [6]
"An evaluation can lie to you convincingly." Microsoft says it has seen perfect scores for code that never compiled, and a "used Platform X" check that passed whether or not the agent used Platform X.
- [7]
The playbook's evaluation model sets out what a result needs before it can be trusted, from criteria that judge meaning to gates that prove the code runs.
- [8]
The playbook shows how to write criteria a judge rules on the same way every time, calibrate them before trusting them, and version them when the product changes, and covers traps such as letting a model write the criteria.
- [9]
The playbook teaches telling apart an extension that never loaded, one that loaded but was never called, and one that was called but applied wrong, because each needs a different fix.
- [10]
The playbook covers testing each proposed change as a hypothesis before shipping it, and bringing evidence to the owning team.
- [11]
The evaluations led to dozens of shipped fixes to docs and agent extensions, including 46 improvements to the Azure Cosmos DB Agent Kit.
- [12]
The SPFx project upgrade runs end to end in the playbook, from the scenario through the fixes that shipped.
- [13]
The method does not depend on a specific evaluation system; teams can use one they built or an existing one, and the playbook lists the capabilities to check for.
- [14]
Along with the PDF, Microsoft is releasing the AX Practitioner skill, which users install in their coding agent and ask questions while working on their own evaluation.
- [15]
Agent-generated code can carry the wrong SDK version, a deprecated authentication pattern, or a setup nobody on the product team would recommend. "The agent didn't make a random mistake. It did exactly what its training data told it to do."
- [16]
"A readout tells you what failed while the trajectory tells you why."
- [17]
Writing criteria is where domain expertise becomes measurable, and it is the hardest step.
- [18]
Nine failure patterns repeat across every technology Microsoft DevRel has evaluated, and each points to the surface to inspect first.
Sources
1 independent publisher whose own reporting we read for this story.
- Introducing the Agent Experience (AX) Practitioner Playbook
devblogs.microsoft.com
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Developer relationsFollow
- Agent Evaluation as Cross-Layer MeasurementFollow
- AI Coding AgentsFollow
- Agent experience (AX)Follow
Entities
- MicrosoftFollow
- AX Practitioner PlaybookFollow
- AX Practitioner skillFollow
- Azure Cosmos DBFollow
- Azure Cosmos DB Agent KitFollow
- SharePoint FrameworkFollow
- Microsoft AzureFollow
- Microsoft 365 CopilotFollow
- Model Context ProtocolFollow