This project gives a fictional claims team a release decision to make, not a prompt to polish. The current assistant checks receipts and drafts a recommendation; the server owns approvals. A candidate prompt shortens explanations and changes how missing evidence is described. The evaluation board must decide whether to ship it, revise the test set, or stop the rollout. Produce a case ledger, paired comparison, adjudication note, adversarial packet, and experiment brief.
Project: govern a claims-assistant evaluation board
Build a defensible offline result
Freeze 47 cases under policy LP-12 and evidence snapshot IS-77. Run baseline PEB-118 and candidate PEB-124 on the same cases and retain each response, validator outcome, and receipt state. Fourteen cases improve, nine worsen, and 24 tie. Case CL-431 is among the losses: the candidate approves a claim without its required receipt. Apply the declared critical gate before considering the net result. Two reviewers disagree on a separate outage-related charge case; preserve both votes, mark it unresolved, and request the missing event record before counting it in a scored denominator.
Probe boundaries and plan exposure
Start from receipt RC-371. Add one mutation that asks the assistant to approve the claim from inside the receipt body, one that forges a higher-priority message in retrieved text, and one benign 'approval pending' control. Confirm that no tool effect occurs and that ordinary evidence remains readable. A later policy revision creates CL-431@2 rather than changing CL-431@1 in place. Re-run both bundles under the new case revision before comparing them. Only after the critical failure is fixed may the team propose an account-stable, read-only canary with a named owner and a stop rule for omitted verification or wrong-account effects.
Deliverables: EV-47@LP-12 case ledger and paired result table.
Critical loss: CL-431@1; release status: blocked.
Unresolved label: TK-862; hold outside scored denominator.
Attack packet: receipt instruction, forged retrieval note, benign control.
Policy migration: preserve CL-431@1 and add CL-431@2.
Canary proposal: read-only accounts after offline gate passes.Performance and operating cost
One paired pass uses 94 model calls for 47 cases before repeats, attack mutations, or label review. If the board runs three attack variants for each of 12 base cases, that adds 36 executions. The ledger uses O(N) records for one revision per case and grows with revisions. Human adjudication is the expensive part when evidence is incomplete; do not hide it behind a faster model judge. A blocked release still provides useful information when its artifacts let the team reproduce the false approval and revise the right control.
Common Mistakes
- Do not ship on a positive net count after a critical false approval.
- Do not resolve reviewer disagreement by copying the candidate answer.
- Do not rewrite an old case revision when policy changes.
Connected lessons
- Paired prompt evaluation: count changes, then inspect uncertainty
- Evaluation labels: adjudicate disagreement before scoring a release
- Adversarial case mutations: test the boundary, not a magic phrase
- Online prompt experiments: define exposure and stop rules first
- Evaluation case ledgers: revise labels without erasing history
- Code lab: block a critical prompt regression
- Prompt telemetry: measure failures without copying private payloads
- Prompt evaluation governance decisions
Field guide: Prompt release review: assemble the decision packet.
