ProvenanceGuard verifies not just facts but their sources

A new paper introduces a verification layer that sits above MCP agents and refuses to accept "the fact exists somewhere in the context" as a sufficient answer. The problem the researchers call cross-source conflation is simple to grasp but hard to catch: the claim is true, the source attributed to it is not. A support agent cites "the account record" when the refund policy lives in a policy document; a clinical agent presents a detail from a patient's history as if it came from the medical literature. In both cases the fact is real, the attribution is wrong, and in a sensitive environment that is no less dangerous than a hallucination.
ProvenanceGuard does not retrain the agent. It runs after the answer has been generated, reads the full trace of tool calls and source IDs, and preserves source identity through the entire pipeline. The process unfolds in five consecutive stages: decompose the answer into atomic claims, route each claim to the most relevant source, run an NLI check to confirm that source actually supports the claim, compare the supporting source against the source the answer cites or implies, and finally issue a verdict at the individual-claim level plus a global allow-or-block decision for the whole answer. Blocked answers can be sent through a RARR-style correction loop and re-verified.
The experimental setup is conservative by design. The tests described in the paper use only local models to maintain a controlled offline environment: MiniLM for source retrieval, DeBERTa NLI for support checking, and a local LLM for claim decomposition. The verifier examines literal values, numbers, dates, and identifiers — it does not settle for a sentence that "sounds plausible." A calibrated decision stage fuses these signals. The conservative policy targets data-sensitive review: better to block an answer than to allow a misattribution, even if that means higher latency.
The named models are the configuration that was tested, not a mandatory requirement of the architecture. The same skeleton — claim, source, decision — can be adapted to managed cloud models, but every new setup will need its own calibration and testing. The results reported in the paper come from the local configuration only, and the researchers stress this explicitly.
Evaluation was carried out on responses produced by a medical agent operating with patient records. The paper is available now on Hugging Face and arXiv as a preprint that has not yet undergone peer review. For teams building agents in environments where misattribution constitutes a regulatory or safety risk, ProvenanceGuard offers a protective layer that does not depend on the agent's own architecture — itself a not-insignificant architectural advantage.