Quality Gates at Every Transition - Stopping a Bad Value Before It Becomes the Next Step's Input
The problem
Take an eight-step insurance claims pipeline. OCR extraction reads the claim form. Field validation checks that the extracted values are well-formed. Policy lookup matches the claim to a record. Coverage determination decides what the policy pays for. A fraud flag screens the claim. Payout calculation sets the dollar amount. An approval letter gets drafted and sent. I ran two hundred synthetic claims through a version of this pipeline with no checks between the steps, a quarter of them carrying a defect I planted on purpose: a smudged OCR amount, a transposed policy number, a coverage edge case. Every step, measured on its own, reported somewhere between ninety and ninety-six percent accuracy. That's the number I'd put on a slide if I were pitching this system, and it wouldn't even be a lie. The OCR model really does read the form correctly more than nine times out of ten. Coverage determination really does apply the right clause almost every time.
Here's what happened to one specific claim. OCR misread a smudged four thousand eight hundred dollars as forty-eight thousand. The digit string was well-formed, so field validation waved it through, because field validation only checks that a number looks like a number, not that it's the right number. Policy lookup matched on the policy number, which the smudge never touched, so that step succeeded too. Coverage determination applied the correct clause to the wrong amount and, by its own standard, succeeded. The fraud flag never fired, because forty-eight thousand dollars isn't unusual for this policy type. Payout calculation computed a check for ten times the correct amount, correctly, given the number it was handed. The letter drafted cleanly and went out. Every single step did its job. The claim was wrong by forty-three thousand two hundred dollars, and nothing caught it, because nothing at any step boundary ever asked whether the output made sense as an input to the next stage.
That's the mechanism I keep running into whenever I look closely at a multi-step pipeline that looked clean in testing. Each step is measured, and passes, against a narrow question: did this step do its own job correctly. The question that actually matters, whether the value crossing the boundary between two steps is sane, never gets measured at all. A pipeline built entirely out of individually passing steps can still produce a confidently wrong final answer, and the wrongness is invisible at every point along the way. It only shows up if someone checks the end, and most pipelines I've looked at don't build anything that checks the end either.
The pattern
The fix is a cheap, narrow check placed at every handoff between steps, run before the receiving step is allowed to see the output. A gate can be simple. It only has to ask one specific, testable question at one specific seam: does the extracted amount fall in a plausible range for this policy type, does the policy number pass a checksum, does the coverage clause the previous step cited actually exist in the policy. When a gate fails, the item's forward progress stops. The item gets routed to a held-for-review queue, and the pipeline moves on without it. A malformed value never gets the chance to become someone else's input.
I ran the same two hundred claims and the same defect rate through a version of the pipeline with a gate after every one of the eight steps, each gate carrying its own stated pass-fail criterion. That one change forces a real shift in how the pipeline has to be built. A pipeline with no gates is a straight sequence: step one feeds step two feeds step three, no branching required. A pipeline with gates has to become something that can branch, because a gate that can fail introduces exactly the logic a straight sequence can't express: pass and continue, fail and route to a held queue, stop this item's forward progress without touching any of the others still moving through the pipeline. I wasn't chasing a specific accuracy target, but I can show that end-to-end success on the identical claims came out materially higher with gates in place than without them, and the tracking I built alongside it shows exactly which boundary catches the most defects, which turns out to matter more than the aggregate number by itself.
What a gate actually protects against is propagation. A step is still going to be wrong some fraction of the time regardless of the gate sitting at its boundary. What the gate changes is whether that wrong output gets the chance to become the foundation for four more decisions built on top of it. Catching a bad value at the boundary where it was produced is cheap. Catching the same bad value three steps later, after it's shaped a coverage decision and a payout calculation, means untangling work that already happened on top of a bad assumption. The gate does the same job a manufacturing line does when it inspects a part at each station instead of waiting for final assembly. Cheap early, expensive late, and the expense compounds with every step the defect survives.
Design considerations
A gate can't exceed the question it was built to ask, and that limit matters more than it sounds like it should. A check that asks whether an amount is plausible will never catch an amount that's wrong but plausible, and a lot of the costliest errors are exactly that: wrong, yet plausible enough to slide under a range check. I read this as guidance for writing a sharper gate. The sharper version queries the actual state of the system it's checking against, ground truth, rather than trusting the producing step's own narrated claim that it succeeded. Reported success and actual success are two different signals, and a gate that reads only the first one is checking the wrong thing.
There's also a question of what decides pass or fail at the gate itself. A hand-written rule works for a lot of boundaries and it's the cheapest thing to build. Training a learned scorer to rank candidate outputs at each step is a real alternative, and it's tempting to treat that as a different pattern. I don't. Swapping a rule for a trained scorer changes only how the check decides. What the gate is for, catching a bad value at the boundary before it becomes the next step's problem, stays the same either way.
The honest limit here is scope. A gate protects only the boundary it sits at. A slow, aggregate decline across many runs, the kind where every individual pass looks clean and the trend is still getting worse, sits outside anything a single-seam check can see, because a gate only ever looks at one item crossing one seam at one moment. It also costs something real: every gate is one more thing to design correctly, and a held-for-review queue that nobody actually reviews is worse than no gate at all, because it creates the appearance of a safety net while catching nothing at all. The gate earns its cost on pipelines long enough, or consequential enough, that a bad value surviving three or four more steps is genuinely expensive to unwind. On a two-step pipeline where a human already reads the final output before it matters, I'd skip it. On an eight-step pipeline that ends in a check going out the door, I wouldn't build it any other way.