Skip to content

Build1 publisher3 min readPublished

MCP and A2A move the work but not the judgment, and audits ask about the judgment

A paper scores five agent interoperability protocols against six governance dimensions and finds they coordinate tasks without expressing who may approve, how dissent survives, or when a human is called.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A paper by Kang and Diponegoro takes five agent interoperability protocols, including MCP and A2A, and scores them against six governance dimensions drawn from organisational theory: membership, deliberation, voting, dissent preservation, human escalation, and audit or replay.
  • The paper's conclusion is that these protocols coordinate tasks but cannot express a governed community, and that governance is a missing architectural layer above these protocols rather than a feature inside them.
  • You cannot state in MCP who is allowed to approve something, how a dissent is preserved, or when a human must be brought in.
  • The Model Context Protocol connects an agent to tools, and A2A lets agents discover one another and exchange messages; both do their jobs well, and neither is trying to do this one.
  • Worked example: an agent drafts a refund decision, a reviewer reads it, changes the amount, adds a line about the customer's contract, and approves. Six months later the agent's draft is in a trace and the final amount is in the payments system, but the reviewer's reasoning was a sentence in a chat thread that has since scrolled away, and the fact that a human changed the number is not recorded anywhere as a distinct event.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A paper by Kang and Diponegoro takes five agent interoperability protocols, including the Model Context Protocol and A2A, and scores them against six governance dimensions borrowed from organisational theory: membership, deliberation, voting, dissent preservation, human escalation, and audit or replay [1]. Their conclusion is that these protocols coordinate tasks but cannot express a governed community, and that governance is a missing architectural layer above them rather than a feature inside them [2].

The gap is not a criticism of either protocol's design. MCP connects an agent to tools and A2A lets agents discover each other and exchange messages, and according to the dev.to post that summarises the paper, both do that job while neither is attempting this one [4]. What you cannot say in MCP is who is allowed to approve something, how a dissent is preserved, or at what point a human must be brought in [3].

The consequence shows up as a missing event rather than a broken call. In the post's worked example, an agent drafts a refund, a reviewer changes the amount, adds a line about the customer's contract, and approves; six months later the draft sits in a trace and the final amount sits in the payments system, but the reviewer's reasoning was a sentence in a chat thread that has scrolled away, and the fact that a human changed the number is not recorded anywhere as a distinct event [5]. The author's position, from building systems for regulated industries, is that a log showing a call was made is not evidence that a person exercised judgment, and clients have to produce that evidence years later to somebody paid to be sceptical [6].

Most implementations start by writing the approval as a boolean [7]. That collapses two different events into one: if the reviewer edited the draft, what went out is not what the agent produced [8]. The proposal is to keep the agent's output as an artefact and the human's intervention as an override carrying the diff, the reviewer's rationale, and a flag for whether the edit refined the agent's intent or replaced it [9]. Refining and substituting are different signals, and a few hundred of them add up to a measure of where the agent fails, produced as a byproduct of the audit trail [10]. Rejections and escalations need the same treatment, since a record of only the decisions that went through keeps the successes and discards the disagreements [11].

Tamper-evidence is the older half of the problem and has a known answer: sign each record and chain it by content hash so each entry commits to the previous one, after which altering any entry breaks every subsequent link and a verifier can re-walk the chain without trusting the system that produced it [12]. For a stronger form, the chain can be anchored in an external transparency log, which is what the IETF's SCITT work addresses; the post's authors offer this as an optional profile and say implementers on the SCITT mailing list corrected several details [13].

The remedy on offer is CHAP, described by its authors as their attempt at this layer, Apache-2.0 licensed and riding on MCP and A2A as transport rather than competing with them [14]. Its Python coordinator has no runtime dependencies, so a clone and one scenario script will show what a decision record looks like [15]. Note that the diagnosis is cited from an independent paper while the remedy comes from the party publishing the argument [1].

Watch whether the SCITT anchoring profile survives implementer review, and whether anyone using this publishes their refine-versus-replace rates, because those numbers are the test of whether the extra record earns its keep [10][13].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories