Assigned to the Case, Not Cleared for the File
Assigned to the Case, Not Cleared for the File
A paralegal gets staffed onto a case. Somewhere in the system that made that assignment, one flag flips from false to true, and from that moment the retrieval layer treats every document tied to that case as fair game. Case correspondence, deposition summaries, scheduling notes, and the handful of privileged attorney work-product memos that were never supposed to leave the attorneys' hands all sit in the same bucket now, because the permission model only ever asked one question: is this person assigned to this case.
The problem
I built a legal case-management retrieval assistant against a corpus of ten case files, six to ten documents each, with a mix of ordinary case correspondence and privileged attorney work product sitting inside the same folders. Eight staff accounts, a blend of paralegals and attorneys, were mapped to cases through a single assignment table. The retrieval check behind the assistant asked exactly one boolean question before returning anything: is this staff member assigned to this case. Say yes, and the whole case file opened up, privileged memos included, because nothing downstream of that check ever asked a second question at the document level.
That single check is coarser than it looks, and the gap it hides has real consequences. There are two different things a person can be entitled to when a case file exists: knowing the case exists and roughly what it contains, and reading the actual content of a specific document inside it. A paralegal assigned to a case is unambiguously entitled to the first. Whether they're entitled to the second depends entirely on the document, and a privileged memo written by an attorney for another attorney's eyes doesn't become paralegal-readable just because the paralegal happens to be staffed on the matter. The case-assignment check answers the first question and the system had been treating it as if it had answered the second one too, which it never did.
I ran twenty queries from paralegal accounts against cases that had at least one privileged document sitting inside them, and measured what I'll call the over-broad exposure rate: how often a returned passage came from a document where the querying user had visibility into the case but no actual access grant to that specific document's content. It came back at twenty percent. One in five of those queries handed a paralegal a passage from a memo they were only ever supposed to know existed, not read. Nobody had to do anything wrong to produce that number. The system did exactly what its single permission check told it to do. The check itself was measuring the wrong thing.
The pattern
The fix is to stop treating case assignment as a proxy for document access and check both grants independently, every time, with neither one allowed to substitute for the other. A visibility check confirms that a user is assigned to the case at all: that grant is what lets them know the case exists, see its metadata, and browse its document list. A separate access check runs against the individual document and asks whether this specific user holds a content-access grant for this specific document, a grant that exists on its own and has nothing to do with whether the case-assignment table says yes. A document only becomes retrievable when both checks clear. Visibility without access still returns something: a metadata stub, enough to confirm a privileged memo exists in the case file, with no path from that stub to what the memo actually says.
The reason I split this into two independent checks rather than one check with an extra condition bolted on is that a single combined check tends to degrade back into the same failure the moment someone touches it under deadline pressure. Split into two named functions, visibility and access, each one is small enough to unit-test in isolation, and you can prove the property that actually matters: that visibility alone, no matter how it's computed, never returns content on its own. On the same twenty queries against the same ten cases, resolving permission at the document level and requiring both grants dropped the over-broad exposure rate from twenty percent to zero, verified against a set of privileged documents spread across multiple cases specifically so no single case's fix could hide a failure elsewhere in the corpus.
The similarity scoring and the ranking stayed exactly as they were through this whole fix. The variable that moved was the granularity permission gets resolved at. Case-level assignment is the coarsest container in this data model, and coarse permission resolution always collapses down to whatever that outermost container allows, because nothing forces the system to look one level deeper before it hands back content. Moving the check to the individual record, the actual document, and refusing to let a coarser grant answer for a finer one, is the entire mechanism: the same filter logic as before, aimed one level deeper than it used to be.
Design considerations
This pattern buys you exactly one property: permission that resolves at the level of the thing actually being returned, instead of at the level of whatever container happens to enclose it. It has real limits past that, and I'd rather name them than let someone assume this closes more than it does.
It doesn't decide who should hold which grant in the first place. Building the two-check machinery tells you nothing about whether the access-grant table itself is correct, whether an attorney who left the case eighteen months ago still shows up with a live grant on a memo they haven't touched since, or whether new privileged documents get their grants set correctly the day they're created. That's a governance and data-hygiene problem sitting upstream of retrieval, and no amount of dual-checking downstream fixes a grant table that was wrong to begin with.
It also doesn't resolve what happens when the two grants conflict with each other for reasons neither one was designed to capture. A litigation hold, for instance, can restrict access to a document independent of both the case assignment and the document-level privilege grant that already exist, and a two-grant model has no built-in way to know a third kind of restriction just showed up. You can bolt a third check on, and at that point you're doing the same thing you did going from one check to two: naming the new condition explicitly rather than letting an existing check quietly try to cover for it.
The metadata stub needs its own deliberate calibration rather than a fixed default. Enough information to confirm a document exists, without leaking anything that makes its content inferable, is a genuinely narrow needle to thread. A stub that includes a document title referencing the substance of a settlement, for instance, can leak the fact you were trying to withhold even while technically respecting the access boundary. I'd rather ship a stub that says less than one that says almost enough, because the failure mode on the wrong side of that line is exactly the leak this whole pattern exists to prevent.
The last thing I'd flag is granularity creep in the other direction. Not every retrieval system needs record-level resolution. If every document inside a container genuinely shares the same access population, adding a second check per document is pure overhead with no property gained. The signal that you actually need this pattern is a real, named exception inside an otherwise-uniform container, a privileged memo sitting among ordinary correspondence, a compensation document sitting among public policy pages. Build the two-grant check where that exception lives. Don't build it everywhere on the assumption that it might.