Order 31 and the Refund That Never Went Through
Order 31 and the Refund That Never Went Through
The problem
I built a small order-support agent to watch a specific failure happen under controlled conditions, because I don't trust an intuition I haven't put through a project of my own. It handles a batch of 40 synthetic order requests, calling three mocked tools in sequence for each one: an inventory lookup, a refund processor, and a shipping tracker. I seeded a handful of those orders with an injected failure, an out-of-stock inventory response here, a refund rejection there, a tracker outage somewhere else, and I deliberately let each mocked tool return its own ad hoc response shape instead of forcing them to agree on one. That's how most tool integrations actually get built, one team writing one tool at a time with no shared contract holding the pieces together.
Order 31 is a refund request. The agent calls the inventory tool, gets back a normal lookup, and moves on to the refund processor. I'd seeded that processor to reject this particular order, but I never told it to reject in any particular shape, so it came back with a nested dictionary carrying a "result": "ok" field two levels down, buried inside a top-level object that, read as a whole, clearly meant the refund had failed. The orchestrator reads only the field that looks like the answer, sees "ok", and waves the order through to the shipping tracker as if the money had already gone back, without ever looking at the object as a whole. Order 31 gets logged as resolved. It wasn't, and nothing in the pipeline so much as flagged it.
Sit with how ordinary that failure is before reaching for a fix. Nothing actually broke. The inventory tool worked. The refund processor worked, in the sense that it did exactly what I'd seeded it to do. The shipping tracker worked. Every individual call succeeded on its own terms, and the customer's refund still didn't happen, because I'd never built the three tools against a shared idea of what "it worked" looks like. Each one invented its own shape for success and failure, and the orchestrator had no way to tell a well-formed lie from a well-formed truth. That's the ordinary condition of any tool catalog nobody designed as a single system, and it's exactly the condition most production agent pipelines are actually running in.
The pattern
The fix is structural rather than clever. Force every tool call through a schema-validated envelope, {status, data, error, tool_name, schema_version}, before the orchestrator or the next tool in a chain reads anything at all. A native response that doesn't map cleanly onto that envelope doesn't get waved through as a courtesy. I treat it as a failure and route it to retry or escalation, which means a mislabeled "ok" field, exactly like the one that sank order 31, can no longer pass as success just because it happened to use the right word somewhere inside it.
I built a second version of the order-support agent to test this against the first one directly. Same 40 orders, same seeded failures, but every tool invocation now passes through a normalizer before the orchestrator sees it. The orchestrator only ever reads the envelope. It never touches a tool's native output again. A response that doesn't fit the envelope's shape gets flagged, the opposite of what happened to order 31 in the first version, where a malformed answer sailed through silently.
What I track across both versions is a silent-failure propagation rate, the percentage of injected failures that go undetected before the next step acts on them, alongside a count of erroneous actions actually taken on bad information, a refund marked complete when it wasn't, a shipment released on an order that should have stopped. I set the bar at a silent-failure rate at or below 2%, with erroneous-action count near zero, across the same 40 orders both versions share. The envelope version is built to catch the order-31 failure mode by construction. The orchestrator no longer has the option of reading a tool's native shape at all, so a buried "ok" field has nowhere to hide.
Design considerations
Here's where the envelope stops, and I think naming the edge matters more than celebrating the fix. The envelope checks whether a response has a well-formed status field. Whether that field's claimed value is actually true is a different question, one the envelope was never built to check. A tool that lies cleanly, one that returns {"status": "success", "data": {...}} for a refund that never processed, sails straight through the envelope without tripping anything, because a well-formed success and a well-formed lie look identical from the envelope's point of view. Closing that gap takes something else entirely: checking a claimed result against an independently observable artifact, a real diff, a real test execution, a state you can verify apart from the tool's own report of itself. The envelope catches order 31's kind of failure, a shape that doesn't parse cleanly, but a shape that parses perfectly and just isn't true sails past it every time.
There's a second edge one layer up that I didn't expect until I went looking for it. The envelope validates structural conformance per call, but it doesn't catch a tool's schema changing shape between calls, a field quietly renamed or retyped that the agent keeps reasoning against without ever crashing. That's schema drift, and it's a genuinely different problem from anything the envelope checks at the moment of a single call. The two failures sit right next to each other in the pipeline, which makes it tempting to assume one fix covers both. What the envelope actually covers is narrower than that: shape conformance at the instant of a single call, full stop.
I'd also think twice before wrapping every tool in a system this heavily. The value of a validated envelope scales with how many independently authored tools are feeding into the same orchestrator, because that's where nobody has agreed on a shape in the first place. A small system built and maintained by one team, where every tool already shares conventions because one person wrote all of them, gets less back for the overhead of a normalization layer on every call. I'd still add the envelope the moment a second team, a third-party MCP server, or anything outside my own direct control enters the picture, because that's the exact point where an assumption about response shape stops being something I can verify by reading the code myself.
The last thing I keep in mind is what happens after a check like this fails. The envelope's whole job is the moment of validation, flagging a response that doesn't fit before anything downstream acts on it. What a validation gate promises is narrower than it sounds: a failure gets noticed, on the record, before it can quietly compound. Deciding whether to retry, escalate to a person, or roll back a partial effect is a separate decision that has to live somewhere else in the system, and those two guarantees only look similar from a distance.