Check It Before It Crosses - Verifying a Claim Before the Next Agent Trusts It

The problem

I built a three-agent news and intel pipeline, and running it unchecked felt like watching a game of telephone with a paycheck attached. A Source agent reads a seeded synthetic corpus of thirty articles, several of them carrying factual errors and outdated figures I'd planted on purpose. A Synthesis agent writes a briefing from whatever Source extracted. A Distribution agent posts that briefing to a shared channel for downstream readers. Nothing checks a single claim against the underlying source text at either handoff, Source to Synthesis or Synthesis to Distribution. It just flows.

Here's what flowing unchecked actually produced. The planted errors survived, unchallenged, all the way into the final briefing. Worse, Synthesis and Distribution sometimes added their own unsupported embellishment on top of what Source had extracted, phrasing that sounded plausible and specific but couldn't be traced back to any actual passage in the source articles, and that embellishment propagated just as unchecked as the original planted errors, because nothing in the pipeline distinguished a grounded claim from an invented one at the moment either kind got written to shared state. I measured claim-survival error rate, the share of planted errors that made it all the way into the distributed briefing, and unsupported-claim rate, new claims that couldn't be matched to any source passage at all, checked against an annotated version of the corpus that marked exactly where each planted error and its correct value lived. Both numbers were bad, and both were bad for the same underlying reason: every agent in the chain was trusting the agent before it by default, with nothing in between doing the trusting for a reason.

That's the part I keep coming back to. Each individual agent in this pipeline is doing its actual job correctly. Source extracts what looks like the relevant information. Synthesis writes a coherent briefing from what it's handed. Distribution posts what it receives. None of them is malfunctioning. The corruption happens specifically at the seams, in the gap between one agent finishing its work and the next agent starting on the assumption that the work it received is trustworthy. A pipeline built entirely out of individually competent agents can still ship a briefing full of errors nobody agent was responsible for catching, because catching them was never anyone's job in the first place.

The pattern

The fix is a verification step that runs before either handoff, checking a candidate claim against its cited source before that claim is written to shared state or handed to the next agent. I inserted a verify-claim node before both the Source-to-Synthesis and the Synthesis-to-Distribution handoffs. It runs a groundedness check, matching a candidate claim against specific spans in the annotated source corpus, and it blocks any claim that fails that match from propagating any further. A claim that can't be traced back to an actual passage in the source material simply doesn't get the chance to become someone else's input.

Running the same thirty-article corpus through this gated version cut both numbers sharply. Claim-survival error rate and unsupported-claim rate both dropped by a wide margin against the unchecked baseline, because planted errors and invented embellishments were caught at the gate before Synthesis or Distribution ever had the chance to act on them. What survived to the final briefing carried a verified-claim rate close to complete coverage, because anything that hadn't been checked against its source simply never made it that far. The mechanism itself is almost boringly simple. No single agent gets smarter. The gate just refuses to let an unverified claim cross the exact boundary where the original failure was happening.

I've found this same shape showing up independently across research I didn't set out looking for, which gives me more confidence in it than any single source would on its own. A claim-verification framework built for checking facts against tabular data runs a verifier step that evaluates an executor's output for consistency before allowing it forward, sending it back for correction rather than letting it propagate. A separate pipeline decomposes and checks a claim before finalizing any verdict on it at all. A third framework, adapting classical methods for verifying the chain of transmission behind a historical account into a formal check for claim-level provenance, validates a serve, review, or quarantine decision at each point a claim moves from one place to another, tested against twenty thousand real claims. None of these are the same system I built. They're the same intent, showing up independently across fact-checking, tabular verification, and provenance research: check it before it moves.

Design considerations

The limit here runs in two directions, and I'd rather name both plainly than let the gate sound more complete than it actually is. First, this gate checks a claim's groundedness against a cited source, which assumes good-faith content that might simply be factually wrong or outdated. It assumes nothing about content deliberately crafted to exploit the fact that a downstream agent treats an inbound message as an authoritative instruction rather than as data worth evaluating on its own terms. A message that's perfectly sourced and verifiably accurate can still carry a goal-hijacking payload, because manipulating an agent's behavior doesn't require a false factual claim at all, and a groundedness gate would pass exactly the content it would need to catch to stop that kind of attack. That's a different threat model than the one this gate was built for, and treating a source-verification check as a security boundary would be a real mistake.

Second, this gate checks a claim at the moment it's written, not whether that same claim is still current by the time a downstream agent actually reads it. Content that was true when Source extracted it and has simply gone stale by the time Distribution posts it sails through a source-citation check untouched, because nothing about staleness makes the original citation wrong. The citation was accurate. The world moved. A gate built to check groundedness against a fixed source has no way to notice that the source itself, or the world it describes, has changed since the check ran.

There's a cost question too, and it's not free to answer honestly. Every claim that gets checked against its source costs something, in latency and in the engineering work of building an annotated reference to check against in the first place. For a pipeline where the source material is stable and well-structured, like the article corpus I built this against, that cost is manageable, and it pays for itself many times over given what it catches. For a pipeline where the underlying source is itself messy, unstructured, or missing altogether, building a reliable groundedness check becomes its own hard problem, and the pattern's value depends entirely on having something solid to check a claim against. A verification gate with nothing trustworthy behind it to verify against is just a checkbox wearing a gate's clothes.