Coherent But Wrong - The Case for a Second Planning Checkpoint
Coherent But Wrong - The Case for a Second Planning Checkpoint
A customer writes in: "Refund to my original card. I already returned the item." I built a support agent to resolve tickets like this against an order database and a policy ruleset: a thirty-day return window, rules distinguishing cash from store credit, exchange eligibility that varies by product category. Given that ticket, the agent plans a resolution. Confirm the return arrived. Confirm it's inside the window. Issue the resolution. Each step follows sensibly from the one before it. I ran a coherence check on that plan, the kind of self-review pass that asks whether the reasoning hangs together, and it passed. Then the agent issued store credit.
The problem
I'm not describing a bug here. I'm describing what happens by default when the only thing checking a plan is a question about whether the plan is internally consistent. Coherence turns out to be an embarrassingly weak bar. It asks whether each step follows from the last, never whether the steps, taken together, land on what was actually asked for or what policy actually entitles the customer to. A plan can pass that bar cleanly and still hand over the wrong resolution, because nothing in a coherence check ever reads the ticket for what the customer explicitly requested, and nothing in it cross-references the specific policy rule that applies to this case rather than the general shape of "refund flow."
I seeded twenty-five tickets against this agent, and I built several of them on purpose to expose the gap: a partial refund where the customer only returned half an order, a category where exchange isn't eligible and a refund is the only lawful path, a payment-method mismatch where the money went out on a gift card but came back requested to a personal one. Coherence-only checking waved most of these through, because a store-credit resolution and a card-refund resolution are both internally coherent plans. They just aren't both correct. When I measured acceptance-criteria satisfaction, meaning the percentage of tickets where the executed resolution actually matched both the customer's stated ask and the applicable policy rule, it landed well below the coherence check's own pass rate. The coherence check was doing its job the entire time. Its job was just never big enough to catch what actually went wrong.
The pattern
The fix is a second checkpoint, not a better first one. I call it the Goal-Validation Checkpoint, and the design choice that makes it work is treating coherence and goal-satisfaction as two separate, sequential questions asked by two separate passes, rather than trying to teach one pass to ask both.
The plan still goes through the coherence check first, the same self-review pass as before. What's new sits after it: a node that independently extracts the ticket's stated acceptance criteria, what the customer actually asked for, and the specific policy rule that applies, then diffs the plan's proposed resolution against both. If the resolution doesn't match, the plan routes back with the mismatch surfaced as concrete feedback, capped at one retry. Two questions, asked in sequence, by two distinct checks that don't share a blind spot.
The reason this has to be two separate passes rather than one smarter pass deserves a closer look, because the intuitive fix is "just ask the same model to think harder." Research into what's been called the self-correction illusion found something specific and uncomfortable about that intuition: a model will reliably catch and fix an error when that error is flagged inside someone else's reasoning, and will reliably fail to act on the identical error when it's sitting inside its own continuous reasoning trace. The paper traces this to how chat templates and role structure work. A model has no learned way to treat a thought-internal substring from its own prior turn as a discrete, actionable object the way it treats a distinct message from another party. Telling a model to double-check its own plan in the same breath it built that plan is asking it to do something it structurally doesn't have the machinery for, no matter how the prompt is worded. That's a structural limitation, and no amount of clever wording closes it. It's why the second checkpoint has to be a genuinely separate pass, complete with its own extraction step and its own comparison.
I built the retry loop narrow on purpose. One re-plan attempt, with the specific discrepancy handed back as the reason. Not an open-ended argument between the plan and the checker, because that just moves the coherence-only failure one layer deeper: a plan and a checker stuck in the same conversation eventually start agreeing with each other for reasons that have nothing to do with correctness.
Design considerations
This pattern is only as good as the thing doing the checking. That limit is real, and I'm not raising it just to sound thorough. If the checker's read of "what the customer asked for" or "what policy requires" can be influenced, gamed, or quietly redefined, either by an adversarial actor or by an agent that's learned it gets rewarded for finding a way past the check, a plan can sail through goal-validation that's been checked against criteria that no longer mean what they're supposed to mean. I haven't found an architecture in the current research that closes that gap cleanly. The pattern assumes the checker itself is trustworthy, and that assumption has to be defended somewhere else in the system.
There's also a cost question that only shows up at scale. Extracting acceptance criteria and the applicable policy rule independently, on every ticket, is real inference spend on top of the plan and the coherence check that already ran. For a support queue processing a few hundred tickets a day, that's a rounding error. For a high-volume, low-stakes flow where being wrong costs almost nothing to fix, the second checkpoint can cost more than the errors it prevents. I'd reserve it for exactly the tickets where a wrong resolution is expensive to unwind: money movement, contractual entitlements, anything where "we'll just issue a correction next time" isn't actually true because there won't be a next time with that customer.
The checkpoint also can't rescue a policy that's genuinely ambiguous. It checks a plan against a stated rule. When the rule itself doesn't clearly cover the case in front of it, a partial return that falls between two documented categories, say, the checkpoint has nothing firm to diff against, and it will either flag everything as a mismatch or let ambiguous cases through inconsistently depending on how the extraction step happens to read the gap. That's a policy-authoring problem wearing a validation-checkpoint costume, and no amount of checkpoint engineering fixes unclear policy.
What convinced me this split matters beyond my own twenty-five tickets is that 2026 agent-evaluation literature keeps landing on the same two-question shape from a completely different angle. Evaluators increasingly treat "plan adherence," whether the agent stayed on the track it laid out, as a distinct measured quantity from "task completion," whether it actually satisfied the goal, specifically because agents can score beautifully on the first and still fail the second. That's the same gap I was staring at when a perfectly coherent plan issued store credit to someone who asked for their card back. Coherent and correct are different questions. A system that only asks the first one will hand you a healthy-looking number measuring entirely the wrong thing.