The problem is a chain of questions.
A real estate site question rarely stops at an address. It becomes a sequence: which parcel is this, what records support its identity, what assumptions drive the economics, which local rules may apply, and what still needs someone to verify?
An AI-generated brief can make that sequence feel complete before the supporting work is complete. Cividian Site Diligence is organized around an investigation pipeline that keeps the supporting evidence and remaining uncertainty visible.
The public repository provides the implementation and a scripted evaluation report. This case study follows those artifacts and the engineering questions they raise.
Keep the stages distinct.
The public workflow progresses from site identity to evidence rows, deterministic financial scenarios, official ordinance discovery, bounded model reasoning, citation validation, and an additional model audit. The output is a ten-section brief with an investigation plan.
The separation matters. A financial calculation can be checked independently of a model’s prose. A source can be retrieved without proving that it applies to a particular parcel. A citation can support a sentence without settling a local authority’s interpretation.
The documented model workflow uses Nemotron 3 Super for bounded reasoning and a separate Nano audit. Official ordinance discovery uses Tavily, with a verbatim quote filter before the downstream reasoning stage.
Three safeguards, three separate jobs
- Validate the output contract. The per-finding validator checks citation identifiers, rejects unverified evidence references, and checks numeric claims against cited rows. A recognized citation or matching number does not by itself establish semantic support, units, or scope.
- Anchor extracted text. The ordinance reader checks that a quote occurs in normalized retrieved text and retains a fetched-text fingerprint. Accepted text remains unverified for parcel applicability.
- Constrain the second reviewer. The auditor can remove or flag unsupported findings, but cannot add a new claim. Malformed verdicts or unsupported spans that do not occur literally in a finding leave the output unaudited.
These safeguards address different failure modes. Their presence makes the mechanism inspectable; a controlled evaluation is still needed to quantify how much each stage helps.
Unknown is a useful output.
Missing values stay null. Zoning remains unverified until confirmation with the planning office. Scanned ordinances and image-based material are outside the documented supported workflow.
A system should expose the boundary between a retrieved fact, a calculated scenario, a model interpretation, and a required external check.
That boundary helps a reader take the next useful action. An incomplete ordinance search should lead to a source request or office confirmation, rather than a more confident paragraph. A scenario with a missing cost input should remain visibly incomplete, rather than acquire an invented estimate.
Geographic coverage also has a boundary: the documented parcel workflow is limited to Indiana without a provider key. A useful interface should show that limitation before a person relies on a result.
Evaluation needs several kinds of evidence.
A fixture can establish that a known input passes through the pipeline and produces the expected structure. It cannot establish that a live ordinance search found every applicable rule, that the source is current, or that a local authority would agree with the interpretation.
The next research step is a frozen, independently reviewed set of site questions with reference sources and explicit unknown states. Results should report source coverage, supported claims, unsupported claims, geographic limits, abstention, latency, and cost separately.
The published evaluation report records 16 scripted pipeline and validator cases. A controlled live benchmark is proposed here to measure source support and real-world coverage.
The larger idea.
For applied AI, the unit of usefulness is often a decision someone can inspect. The answer needs a source trail, a clear account of its assumptions, and a concrete path from uncertainty to further investigation.
Cividian Site Diligence is a public example of that direction: separate calculation from interpretation, make citations inspectable, and preserve what the system does not know.
Architecture and limitations are drawn from the public project README at a pinned revision. The repository is published under Apache-2.0. Interpretive commentary and proposed evaluation work are presented here as self-published engineering analysis.
Have a question or a useful counterexample?
Start a conversation