An answer is easier to trust when a reader can see how it was built. That requires more than a list of links at the bottom.
Start with the source trail.
When a system retrieves documents and then produces prose, it crosses several boundaries. A document can be relevant without being authoritative. It can be authoritative without being current. It can be current without applying to the specific situation in front of the reader.
A source trail should preserve the document’s identity, the retrieved material, its date context, and which claim it supports. The interface should let a person inspect these relationships without reverse-engineering the model’s reasoning.
In the public Cividian Site Diligence workflow, evidence rows precede reasoning, ordinance discovery includes a verbatim quote filter, and citation validation is a separate stage. That architecture is a useful starting point for treating evidence as an object the product carries through the workflow.
Different claims need different labels.
Consider four statements in a site investigation:
- Retrieved fact: a record identifies a parcel or quotes a rule.
- Calculated scenario: an output follows from explicit financial inputs and formulas.
- Interpretation: the system reasons about what a source could mean in context.
- Open question: the available evidence does not settle the answer.
These statements have different failure modes. A retrieved fact can be linked to the wrong site. A calculation can use a wrong assumption. An interpretation can exceed its source. An open question can disappear into smooth, plausible prose.
Keep the evidence status close to the claim it qualifies.
A generic confidence number is less useful than a specific boundary: “missing input,” “source date unknown,” “interpretation requires confirmation,” or “contradicting sources.” These labels tell the reader what to investigate next.
Unknowns should survive the pipeline.
If an input is unavailable, preserving null keeps the absence visible to downstream code and the reader. Replacing it with a plausible value makes an incomplete calculation look finished.
The same principle applies to prose. A system should be able to say that no supporting source was retrieved, that a source cannot be read, or that a local office must confirm an interpretation. An abstention can be the correct result of an otherwise successful workflow.
Measure the interface as well as the answer.
A useful evaluation would ask reviewers to identify which claims are supported, which assumptions affect a calculation, and which questions remain unresolved. Compare ordinary answer text with an interface that exposes evidence status, using the same questions and source set.
This is a proposed experiment, not a reported result. The hypothesis is that an inspectable evidence interface helps readers notice unsupported claims and choose better next actions. It could also add friction, and the evaluation should measure that cost.
The architectural example comes from the public project README at a pinned revision. This note is self-published engineering commentary, not a peer-reviewed paper.
Have a question or a useful counterexample?
Start a conversation