Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Microsoft DevRel's AX playbook traces coding-agent mistakes back to docs, skills and MCP tools

Microsoft DevRel published the AX Practitioner Playbook, a method for finding why coding agents misuse a platform, drawn from hundreds of agent sessions. Its case is that owners of docs, MCP servers and CLIs can fix agent behavior now, if their evals can prove which change helped.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Microsoft DevRel's AX playbook traces coding-agent mistakes back to docs, skills and MCP tools
Generated illustration

What happened

  • Microsoft DevRel has measured coding agents on Azure, Cosmos DB, SharePoint Framework and Microsoft 365 Copilot extensions since fall 2025, using prompts like those real developers write.
  • The post names docs, MCP tools, skills, plugins, instructions, CLIs and APIs as the agent sources a platform team can actually change.
  • The playbook walks through an SPFx project upgrade end to end, from the test scenario through the fixes that shipped.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Teams already scoring agents with model judges may be reporting passes for code that never compiled, unless a gate actually executes the output.
  • decision Before rewriting docs, an SDK or MCP server owner has to establish whether the agent loaded and called their extension, since a content fix cannot reach a skill the agent never opened.
  • cost The expensive input is domain experts' time writing, calibrating and versioning criteria, work the playbook says should not be handed to a model.

"An evaluation can lie to you convincingly," Microsoft's post says, and its examples are specific [6]. It reports perfect scores for code that never compiled, and a "used Platform X" check that passed whether or not the agent used Platform X [6]. A check that passes on every input measures nothing. The playbook's evaluation model runs from criteria that judge meaning to gates that prove the code runs [7]. The run gate is the cheapest layer to build, and it would have caught the first of those two failures.

Microsoft calls writing criteria the hardest step [17]. The playbook asks for criteria a judge rules on the same way every time, calibrated before anyone trusts them and versioned when the product changes [8]. It warns against letting a model write them [8]. I think the ordering is correct. A judge that flips its verdict on a rerun cannot tell you whether last week's docs edit helped, and the playbook wants every proposed change tested as a hypothesis before it ships [10].

"A readout tells you what failed while the trajectory tells you why," Microsoft wrote [16]. The split I would adopt first sorts failures into three cases, since each calls for its own fix: the agent never loaded the extension, loaded it but never called it, or called it and applied it incorrectly [9]. In my view the first two are discovery problems. Rewriting the content of a skill the agent never opened changes nothing. The playbook also catalogs nine failure patterns that Microsoft says recur across every technology it evaluated, each pointing to the surface to inspect first [18].

The post's account of a bad generation is blunt: the wrong SDK version, a deprecated authentication pattern, and an agent that "did exactly what its training data told it to do" [15]. It adds that a knowledge cutoff "tells you little about what a model knows about your product" [3]. If the cutoff is no guide, the prior has to be measured product by product. Every lever the post names [4] changes behavior only if the agent reaches it during the session. So the loaded-or-called question comes before any content fix.

The outcome evidence is thinner. Microsoft credits the evaluations with dozens of shipped fixes, including 46 improvements to the Azure Cosmos DB Agent Kit [11]. Forty-six is a count of changes. The post does not report how agent pass rates moved after them. The scenarios ran on Microsoft's own products [5]. For the results to carry to another platform, its developers' prompts have to resemble the scenarios its team writes, and its agents need a route to its extensions at all. The method does not require a particular evaluation system, and the playbook lists the capabilities to check for in whichever one a team uses [13]. Most of the adoption cost is therefore expert time spent on criteria.

Microsoft also released the method as an AX Practitioner skill to install in a coding agent and question during an evaluation [14]. Like any skill, it has to be loaded before it can help.

What to watch

  • Whether Microsoft publishes before-and-after agent pass rates for the Cosmos DB Agent Kit or SPFx scenarios, which would turn a count of fixes into an outcome.
  • Whether the nine failure patterns hold when teams outside Microsoft run the method on non-Microsoft SDKs and MCP servers.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence40
Adoption15
Hype gap+10
Incentives55
Confidence45
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Microsoft DevRel published the AX Practitioner Playbook, described as its method for evaluating what AI coding agents do with a technology, diagnosing why they get it wrong, and fixing it at the source; it is downloadable at aka.ms/ax-playbook.

    ReportedSupportedSource: Microsoft DevBlogs postView cited source
  2. [2]

    The playbook puts in one place what Microsoft DevRel learned from hundreds of agent sessions.

    ReportedSupportedSource: Microsoft DevBlogs postView cited source
  3. [3]

    "Waiting for models to get better isn't a strategy: a knowledge cutoff tells you little about what a model knows about your product."

    ReportedSupportedSource: Microsoft DevBlogs postView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. devblogs.microsoft.com

    1 article · October 9, 2026

    Introducing the Agent Experience (AX) Practitioner Playbook

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Entities

Loading related stories