Build a small internal policy answer system that keeps revisions and permissions intact, measures retrieval, and abstains when evidence is missing.
Project: answer support-policy questions with permission-aware evidence
Prepare the corpus
Create a fixture of current, retired and restricted handbook pages. Include two policies with similar headings and one exception sentence. Record document ID, revision, effective window and allowed department scope. Chunk without separating the exception from its rule. The corpus contract] and chunk identity] are required deliverables.
Build and judge retrieval
Implement a lexical baseline and a second retrieval signal. Create judged questions for exact code, paraphrase, two-section and no-answer cases. Report required-passage recall at the context budget, not only an end-to-end answer score. A restricted page should never appear for an unauthorized principal. Layered evaluation] guides the report.
Constrain the answer
Return an answer state and internal evidence IDs. Reject a claim without an available supporting passage and abstain on stale or conflicting policy. Add a hostile passage that asks for another department’s content; application-level permission and tool checks must still hold. Trust boundaries] are a project acceptance test.
Prove update behavior
Replace one current policy and delete another. Build a candidate index, deliberately fail one chunk, and show that the old complete generation remains active while revoked access is blocked. Submit fixture, scripts, judged set, metric counts, error ledger, latency budget and failure-case output. A fluent demo answer alone does not complete the exercise.
Implementation
def check_policy_project(answer, retrieved_ids, principal_allowed_ids):
evidence_ids = set(answer["evidence_ids"])
if not evidence_ids.issubset(set(retrieved_ids)):
raise ValueError("answer names unseen evidence")
if not evidence_ids.issubset(set(principal_allowed_ids)):
raise PermissionError("answer exposes restricted evidence")
if answer["status"] == "answered" and not evidence_ids:
raise ValueError("unsupported answer")
return TruePerformance and operating cost
Index construction scales with corpus size and embedding work; each query adds retrieval, optional reranking and generation latency. Keep p95 by stage and temporary storage for index generations in the project report.
Common Mistakes
- Do not count restricted passages as valid retrieval hits.
- Do not accept an answer with evidence IDs absent from its context.
- Do not publish a partly rebuilt index.
