A test passing is evidence about a particular check. The useful question is what that check actually established.
There are several layers of correctness.
In an applied AI workflow, software correctness, source coverage, claim support, and real-world applicability are related but distinct. A system can pass its software checks while producing an incomplete investigation.
Did the mechanism behave as expected?
Known inputs, calculation invariants, output structure, failure handling, and citation formatting.
Does the result have the support it needs?
Source identity, currentness, coverage, claim support, contradictions, and unresolved questions.
A third layer is external confirmation. In a site investigation, a planning office’s interpretation may be required even when the source retrieval and software behavior are sound. That confirmation belongs to a different evidence process.
Fixtures make behavior reproducible.
A fixture is a controlled input and expected result. It is especially useful when the live environment changes: documents move, search results vary, APIs fail, and models can produce different outputs.
Fixtures can test a deterministic calculation, verify that missing values stay missing, or ensure that a citation validator rejects an unsupported reference. They reduce uncertainty about the mechanism. They do not reproduce all uncertainty in the world.
The public Cividian Site Diligence documentation distinguishes fixture behavior from live quality. Its architecture also separates deterministic financial scenarios from reasoning and audit stages. The distinction should remain visible whenever results are reported.
A live evaluation needs a reference.
A convincing live output is an example, not an accuracy estimate. A useful evaluation needs a defined question set, independently reviewed reference material, a consistent grading rubric, and a record of the system configuration.
For a civic-document workflow, the test set should include missing evidence, conflicting rules, superseded documents, and material the pipeline cannot read. Easy, well-documented cases tell only part of the story.
Report source coverage and unsupported claims separately. Report abstention separately from incorrect answers. Keep jurisdiction, document date, latency, and cost attached to the result so another person can understand its scope.
Keep the receipt attached to the claim.
A reproducible result should identify the source revision, inputs, environment, configuration, and evaluation method. Without that record, “it worked” is difficult to inspect or repeat.
State the check that ran, the result it produced, and the boundary of the conclusion.
That habit makes progress easier to assess. It also makes failure more useful: a broken stage, missing document, or unsupported claim becomes a concrete engineering question.
The project example comes from the public project README at a pinned revision. The evaluation practices described here are proposals and engineering principles; this note does not report a completed benchmark or an independently validated outcome.
Have a question or a useful counterexample?
Start a conversation