Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: verify a contract-renewal prompt at the boundary

Last updated: 5 Oct 202618 min read
project
AdvancedBy AITrove Editorial

This project tests a contract-renewal assistant against a fixed policy boundary. The fictional service manages industrial equipment, and renewal RN-284 depends on an inspection no older than 47 days plus a clear safety-hold status. The deliverable is a prompt, a compact evidence packet, a case set, and a release report. No answer is allowed to become an external renewal effect until the rule checks pass.

Prepare the evidence

Index policy POL-64 version 4, including clause 8 and appendix C. Retain the inspection record, the hold record, their asset IDs, and their dates. Compress the packet only after checking that both the main clause and exception survive. Write three contrastive demonstrations: eligible at 46 days, ineligible at 48 days, and review with unknown hold status. Keep the 47-day boundary and unrelated asset records for the holdout rather than copying them into the prompt.

Run and compare

Build 43 test cases and hold out 14. Include a stale policy, absent inspection, mismatched asset ID, malicious maintenance note, and conflicting hold versions. Ask for decision, rule ID, evidence IDs, and a short rationale. Generate three candidates only for cases that fail the single-candidate baseline. Filter candidates by schema and current evidence IDs before ranking. Send disputed claim support to a reviewer; a majority vote cannot override a missing record. Allow one critique revision when it names a concrete defect and rerun all gates afterward.

Output
Policy: POL-64 v4, clause 8 and appendix C.
Case: RN-284, inspection age 47 days, hold status clear.
Boundary tests: 46, 47, 48 days; missing hold; other asset ID.
Output: decision, rule_id, evidence_ids, concise rationale, unknowns.
Gate: no approval without both current conditions; no effect from negative controls.
Report: 43 cases, 14 held out; wrong decisions and withheld share by slice.

Performance and operating cost

A single model call per case is the baseline. Three candidates on only failed or disputed slices add cost where it has a plausible benefit, rather than tripling every request. Context distillation adds an initial pass but can save repeated tokens if policy versions remain pinned. Record model calls, output tokens, validator time, reviewer minutes, and forbidden effects. Stop if the multi-candidate path does not improve the held-out failure classes after its extra cost.

Common Mistakes

  • Do not accept the most common answer when all candidates cite stale policy.
  • Do not drop the safety-hold exception from the compact packet.
  • Do not let a polished rationale bypass the date and hold checks.

Connected lessons

prompt engineering
evaluation
Storage details