Skip to content

Build1 publisher3 min readPublished

ProvenanceGuard flags MCP agent answers that credit a true fact to the wrong tool

Multiverse Computing's ProvenanceGuard caught 138 of 139 bad claims experts flagged in a medical MCP agent by checking which tool each fact came from. A source-blind check passes a true fact credited to the wrong tool, so the verifier needs traces that keep each output tagged with its source.

The Engineer · Build desk

Illustration accompanying ProvenanceGuard flags MCP agent answers that credit a true fact to the wrong tool

What happened

  • The target failure is cross-source conflation, where a claim is true somewhere in the pooled evidence but credited to the wrong source, so a source-blind verifier can pass it.
  • It splits an answer into claims, routes each to a source, scores support, compares that source with the one the answer names or implies, then allows or blocks.
  • Blocked answers can go through a RARR-style repair or a safe fallback, and the verifier checks the revised answer again.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams that gate MCP agents on faithfulness scores alone now have to decide whether attribution gets its own pass-fail check, because pooled scoring cannot see a misattributed claim.
  • constraint Agents that merge tool results into one untagged context cannot use this approach until their tracing records a source ID for every tool output.
  • cost The 138-of-139 result belongs to the local MiniLM and DeBERTa setup, so a team moving to hosted models pays for its own testing and calibration before trusting the number.

In the authors' support-desk example, an agent answers, "According to the account record, this plan includes a 30-day refund window." The window is real, but it appears in a policy document [3]. Pool the tool outputs into one context and a support check finds the fact, so the claim passes [2]. ProvenanceGuard never pools them [5]. The claim routes to the policy document and scores as supported. It then fails the comparison against the account record the answer named [6]. The clinical version is a medication detail from a patient-history tool, presented as a finding from the medical literature [4]. The authors wrote that "in a data-sensitive setting a wrong attribution can be as damaging as a wrong fact" [16].

The evaluated build runs on small local models. MiniLM finds the relevant source, a DeBERTa NLI model judges support, and a local language model splits the answer into claims [7]. A calibrated decision step combines those signals [7]. The piece I would copy first is the literal-value rule. A number, date, or identifier absent from the source cannot pass because the sentence sounds plausible [8]. Repairs get the same treatment. A rewrite can introduce a fresh misattribution, so a blocked answer's RARR-style revision or safe fallback goes back through the verifier [9].

The design rests on the trace. The layer reads captured MCP tool outputs with their source IDs and needs no retraining of the agent [5]. According to the authors, the method applies outside medicine when an agent keeps that record [15]. An agent that concatenates tool results into one untagged prompt leaves this verifier nothing to compare. Implied attribution is the softer spot. Answers carry provenance explicitly, as in "according to the account record," and sometimes only implicitly [17]. The attribution step compares against whichever source the answer names or implies [6].

The test data came from a medical agent that used patient records, research articles, and other tools, giving 281 real traces [12]. The authors chose medicine because a patient's record and general research cannot be treated as the same source [18]. Experts labeled 361 claims from 40 answers held out from development [13]. That is about nine claims per answer [4]. They marked 139 as should-not-pass, and ProvenanceGuard caught 138 of them [14]. Recall on the flagged set is 99.3% [1]. The flagged share was high: 139 of 361, or 38.5% [2]. The recall figure alone does not show how many of the other 222 claims were blocked [3]. The authors call the decision policy conservative and say it suits data-sensitive review, where getting the source right matters more than speed [11].

For the recall to carry over to another agent, that agent has to log a source ID with every tool output. Its sources also have to be about as separable as a patient chart and a journal article. The authors say the named models are not a requirement. They say a hosted-model setup would need its own testing and calibration, and that the reported results come from the local configuration [10]. In my view the separation logic transfers now, and the 99.3% transfers only after a team recalibrates on its own traces [1][10].

What to watch

  • The block rate on the 222 claims experts judged passable, which the 138-of-139 recall figure does not cover.
  • A published result from a hosted-model configuration or a non-medical domain, which the authors say needs its own calibration.
  • Whether agent frameworks log a source ID with each tool output, the precondition the authors set for use outside medicine.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories