Why an AI Finding Without a Citation Is Worthless

An evidence-gated finding is an AI-generated finding that is only permitted to reach a human reviewer if it carries a verifiable citation — document, revision, page, and the specific text or drawing region it relies on — and that citation has been independently re-checked against the source. Everything else is quarantined. This sounds like pedantry until you understand the alternative: without the gate, an AI review tool is a machine for generating confident claims at industrial speed, some unknown fraction of which are false, delivered in a tone that makes the false ones indistinguishable from the true ones.
The problem the gate solves
Language models hallucinate — they produce fluent, well-formatted assertions unsupported by their sources. This is a structural property of how they work, not a defect any vendor has patched out (we cover the failure modes fully in the honest guide to AI document review). In construction review the consequences are specific and nasty:
- A finding asserting the spec requires a slip rating the spec never mentions — and a subcontractor priced against it.
- A "conflict" between a drawing and a schedule where the model misread one side of the comparison.
- A reference to an Australian Standard requirement that paraphrases the standard into something it doesn't say.
The asymmetry is what kills you: a false finding costs an hour to disprove, but an undetected false finding that flows into an RFI, a variation assessment or a compliance report costs credibility — and in a review context, credibility is the entire product. One fabricated citation discovered by a superintendent and every future finding from the tool gets re-checked from scratch, at which point the tool has negative value.
What a real citation looks like
Not "per the specification" — that's decoration. A citation that earns the name is resolvable in one click to the exact evidence:
- Document + revision. A-951 Rev D, not "the door schedule" — because on a live project there are four door schedule revisions and the finding is only true against one of them.
- Page and location. The page number and, for drawings, the region — a bounding box the viewer can highlight, so the reviewer's eye lands on the detail, not the sheet.
- The stated content, quoted. What the document actually says at that location, so the claim and its basis sit side by side and any daylight between them is visible.
- Both sides, for conflicts. A cross-document finding ("drawing states X, schedule states Y") needs the full citation for each side. Half-cited conflicts are the most common way a misreading hides.
This is the standard your own profession already applies. A well-formed RFI cites document, revision, page and the conflicting statements — because uncited questions don't get answered and uncited claims don't get believed. The gate simply refuses to let the machine work to a lower standard than you do.
The gate, mechanically
Three parts, all architectural — meaning the system enforces them, rather than trusting anyone's discipline:
- Generation must emit citations. Every candidate finding is produced with its claimed sources attached. No citation, no candidacy.
- Verification is a separate pass. An independent step re-opens each cited location and checks two things: the location exists (right document, right revision, real page) and its content actually supports the claim. This matters because generation grading its own output is exactly the failure you're defending against — the verifier must be a different process reading the real source, not the model agreeing with itself.
- Failures are quarantined, visibly. Output that fails verification or falls below a confidence threshold goes to a quarantine tray — inspectable, but never rendered as a finding. Visibility matters twice over: the reviewer can audit what was held back, and the tray's volume is diagnostic. A quarantine tray that suddenly fills on one document set usually means bad scans or hostile table layouts — a reading problem worth knowing about before you trust the findings that did pass.
Checkable is not the same as correct
Honesty requires the caveat: the gate makes findings defensible, not infallible. A citation can resolve to a real page and the interpretation can still be wrong; a finding can be textually perfect and practically stale because the superintendent resolved the conflict verbally last Tuesday. The gate's job is narrower and more important — it guarantees that every claim put in front of you can be checked in seconds, at the evidence, so that human adjudication is fast enough to actually happen. That's the economics of the whole arrangement: 200 cited candidate findings can be triaged in a morning because each opens onto its proof; 200 uncited claims would each need independent research, and nobody does 200 research tasks — they either trust blindly or abandon the tool.
Which is why evidence gating and human sign-off are one system, not two features. The gate makes the machine's output checkable; the human check makes it decidable; the audit trail makes the decision defensible when someone later asks how it was reached. In ParitySense every confirmed finding carries its citation into the evidence viewer, so the chain from claim to page to human decision never breaks.
FAQ
What is an evidence gate in AI document review?
An architectural rule that no AI-generated finding reaches the reviewer unless it carries a verifiable citation — document, revision, page, and the specific text or drawing region — that has been independently re-checked against the source. Output that fails the gate is quarantined, not shown as a finding.
What is a quarantine tray?
A visible holding area for AI output that failed citation verification or fell below a confidence threshold. It preserves the output for optional inspection without letting it masquerade as a finding — and its volume is a useful signal of how well the system is reading a given document set.
Doesn't a human reviewer checking everything make the gate redundant?
No — the gate is what makes human review feasible. Two hundred cited claims can be adjudicated in a morning because each opens onto its evidence; two hundred uncited claims would each need independent research, which at scale nobody does.
Is a citation enough to make a finding correct?
No. A citation makes a finding checkable, not correct. The cited page can be real and the interpretation still wrong, or overtaken by project context the documents don't contain. That is why evidence gating pairs with human adjudication rather than replacing it.
Run it through ParitySense alongside your manual review and measure the delta.
See how it works