Retrieved Correctly, Quoted Faithfully, Still Wrong
Every layer of verification we have built asks whether the answer matches the source. None of them asks whether the source is right, and for a corpus assembled over two decades that is not a theoretical gap.
The Value That Was Always There
A maintenance interval given as twelve months in a procedure document, and as six months in the equipment specification. Both documents are in the corpus, both are current, and the procedure has been wrong since a revision in 2019 that nobody caught.
The assistant answered twelve months, quoting the procedure, with a link to the section. Our quotation check passed, our reconciliation check had nothing to reconcile against because it operates within a record rather than across documents, and the answer was wrong.
What Our Checks Actually Verify
Faithfulness. Does the answer say what the source says. That is the property most retrieval systems verify and it is the one worth verifying, because the alternative failure, an answer with no source, is more common and easier to produce.
It is not correctness. A faithful answer from a wrong document is indistinguishable from a faithful answer from a right one by any check that looks only at the pair, and the entire industry vocabulary around grounding describes faithfulness while sounding like it describes truth.
The Research on Numbers Specifically
Patil published work in 2026 on detecting manipulated numerical claims in retrieval systems used in a government setting, treating figures as a category deserving their own verification rather than as ordinary text.
The security framing is not ours, since nobody manipulated our document. The useful part is the category: numbers are where a wrong source does the most damage, they are extractable and comparable in ways prose is not, and they can therefore be checked across documents where prose cannot.
| Check | What it can establish |
|---|---|
| Quotation exists in the source | Faithfulness. Not correctness |
| Values reconcile within a record | Internal consistency |
| Same quantity across documents | Contradictions in the corpus |
| Value against a system of record | Correctness, where one exists |
What We Built
A cross-document consistency pass over extracted numeric claims. For each piece of equipment, values with the same name and unit are compared across documents, and disagreements are listed for the customer's technical documentation team rather than resolved by us.
The first run produced ninety-one disagreements across the corpus, of which the team judged sixty-two to be real errors in one document or the other. Nobody had known, and several had been repeated in answers for as long as the assistant had existed.
What the Assistant Does Now
Where a value has a known disagreement, the answer says so and gives both, with their sources and dates. It does not choose. Choosing would require judgement about which document governs, which belongs to the customer and in one case took them three weeks to establish.
That is a worse user experience and the correct one. An assistant that picks a value silently is confidently wrong half the time in exactly the cases where being wrong matters, and users told us that seeing the conflict was more useful than an answer would have been.
Why Prose Cannot Be Checked This Way
Two documents describing a procedure differently are usually not in conflict; they are written for different audiences or different variants. Extracting a claim from prose and comparing it across documents produces mostly false positives, and we tried.
Numbers with names and units are the exception. Twelve months and six months for the same interval on the same equipment is a contradiction with no benign reading, and that narrowness is why the check works at all.
The Part That Stays a Data Problem
We can surface contradictions. We cannot fix documents, decide which is authoritative, or prevent the next wrong revision, and it would be inappropriate for us to do any of those on a customer's technical documentation.
What changed is that the corpus now has a defect list with owners. That is unglamorous document management rather than anything to do with models, and it is the largest quality improvement this system has had in a year.
What We Do Not Claim
We do not claim we detect wrong sources. We detect disagreements between sources. A value that is wrong in every document that mentions it, or that appears only once, passes every check we have and always will.
We also do not claim the sixty-two figure says something about this customer. A corpus assembled by many people over twenty years contains contradictions, and the surprising thing is not that they exist but that nothing had ever looked.
