This project tests a contract-renewal assistant against a fixed policy boundary. The fictional service manages industrial equipment, and renewal RN-284 depends on an inspection no older than 47 days plus a clear safety-hold status. The deliverable is a prompt, a compact evidence packet, a case set, and a release report. No answer is allowed to become an external renewal effect until the rule checks pass.
Project: verify a contract-renewal prompt at the boundary
Prepare the evidence
Index policy POL-64 version 4, including clause 8 and appendix C. Retain the inspection record, the hold record, their asset IDs, and their dates. Compress the packet only after checking that both the main clause and exception survive. Write three contrastive demonstrations: eligible at 46 days, ineligible at 48 days, and review with unknown hold status. Keep the 47-day boundary and unrelated asset records for the holdout rather than copying them into the prompt.
Run and compare
Build 43 test cases and hold out 14. Include a stale policy, absent inspection, mismatched asset ID, malicious maintenance note, and conflicting hold versions. Ask for decision, rule ID, evidence IDs, and a short rationale. Generate three candidates only for cases that fail the single-candidate baseline. Filter candidates by schema and current evidence IDs before ranking. Send disputed claim support to a reviewer; a majority vote cannot override a missing record. Allow one critique revision when it names a concrete defect and rerun all gates afterward.
Policy: POL-64 v4, clause 8 and appendix C.
Case: RN-284, inspection age 47 days, hold status clear.
Boundary tests: 46, 47, 48 days; missing hold; other asset ID.
Output: decision, rule_id, evidence_ids, concise rationale, unknowns.
Gate: no approval without both current conditions; no effect from negative controls.
Report: 43 cases, 14 held out; wrong decisions and withheld share by slice.Performance and operating cost
A single model call per case is the baseline. Three candidates on only failed or disputed slices add cost where it has a plausible benefit, rather than tripling every request. Context distillation adds an initial pass but can save repeated tokens if policy versions remain pinned. Record model calls, output tokens, validator time, reviewer minutes, and forbidden effects. Stop if the multi-candidate path does not improve the held-out failure classes after its extra cost.
Common Mistakes
- Do not accept the most common answer when all candidates cite stale policy.
- Do not drop the safety-hold exception from the compact packet.
- Do not let a polished rationale bypass the date and hold checks.
Connected lessons
- Reasoning summaries: show checkable grounds, not invented certainty
- Multiple candidates: filter invalid answers before ranking
- Critique loops: require a named defect and a stopping rule
- Few-shot examples: teach the boundary with near misses
- Context distillation: shorten input without losing governing exceptions
- Negative controls: test the answer that should not be produced
- Code lab: reject unknown evidence IDs
- Code lab: block a critical prompt regression
- Human handoff: preserve evidence and the reason for uncertainty
- Reasoning and evidence checks
