This project measures a retrieval-backed warranty assistant before it can make a customer recommendation. Assemble 61 fictional questions against versioned plan documents. Include short direct questions, multi-turn references, outdated clauses, and cases where a rule and its exception fall in different chunks. The deliverable is a retrieval and answer-quality packet, not a single polished answer.
Project: measure a retrieval-backed answer gate
Prepare the evidence set
For each question, record the governing plan ID, effective version, required passages, and an expected answer or review state. Reserve 19 questions as a holdout. Run the original query and a self-contained rewrite, recording both. For each retrieval strategy, measure whether all required passages appear in the top five results. A passage that mentions the product but omits the exception does not count as sufficient. Expand parent context only inside the authorized document scope.
Gate the decision
Ask the model for decision, evidence IDs, and uncertainty reason. Validate that cited IDs are in the supplied set and that the policy version covers the case date. A deterministic code path checks arithmetic for any price adjustment. The answer gate withholds a final decision when evidence is missing, conflicting, or expired. Then validate the response again for its destination: render customer text as text and run any proposed account effect through the executor's permission checks.
Question set: 61 cases; 19 untouched holdout cases.
Retrieval gate: all required passages in top 5 and current version.
Answer gate: decision supported by supplied IDs; otherwise review.
Compare: original query vs verified rewrite, same cases.
Report: complete-evidence rate, unsupported-answer count, withheld share.Performance and operating cost
For N questions and Q query variants, retrieval runs are O(NQ) before any model repeats. Parent expansion increases context tokens, so measure cost per accepted case as well as evidence completeness. Keep retrieval errors separate from answer errors and publish raw counts for critical failures. A candidate that answers fewer cases may still be preferable for a consequential action if it avoids unsupported decisions, but the added human-review load must be explicit.
Common Mistakes
- Do not count a topical hit as the full governing rule.
- Do not tune query rewriting on the untouched holdout.
- Do not let a model-generated decision bypass the output sink or executor.
Connected lessons
- Retrieval queries: resolve references without adding facts
- Document chunks: preserve the clause and its governing exception
- Selective answers: measure when to abstain
- Evaluation leakage: keep the release test independent
- Generated output: validate again at the destination boundary
- Prompt evidence and output decisions
