Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Evaluation case ledgers: revise labels without erasing history

Last updated: 5 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

An evaluation case is more than a prompt and preferred sentence. It includes the legitimate task, input snapshot, evidence IDs, policy version, expected decision, permitted actions, severity, and label status. A ledger records each revision and why it happened. When a policy changes, create a new case revision or an explicit migration; do not overwrite the old expected outcome and then compare new scores with an earlier run as if the test were unchanged. Mark cases used for prompt tuning separately from held-out cases, and control access to the holdout. Keep private payloads out of ordinary fixture exports.

Decision in practice

Case CL-431 originally required review when a receipt was missing under policy LP-12. A later policy allows low-value claims to proceed with another verified document. The board creates CL-431 revision 2 with the new evidence rule and a linked policy decision. Old release results continue to point to revision 1. A candidate tested on revision 2 cannot be compared directly with a baseline run against revision 1. The board either replays both bundles on the new set or reports the scores as separate regimes. The older case remains useful for reproducing an incident from its release period.

Output
Case ID: CL-431; revision: 2; policy: LP-13.
Input snapshot: IS-77; evidence IDs: RC-371, VD-9.
Expected decision: review unless VD-9 is verified.
Severity: critical; action permission: none.
Label state: adjudicated; prior revision: CL-431@1.
Use: holdout; access: evaluation owner only.

Performance and operating cost

A ledger lookup can be O(1) by case ID and revision; storing R revisions across N cases takes O(NR) records in the worst case. Replaying both bundles after a policy migration costs about 2N model calls for one pass, but it is the price of a valid comparison. Snapshot retention and access control may dominate storage cost for document-heavy cases. Keep metadata and evidence references when full payload retention is not necessary. A test set that silently changes labels makes release trends meaningless even if every individual score was computed correctly.

Common Mistakes

  • Do not overwrite a historical label in place.
  • Do not compare scores across different policy revisions without replaying both sides.
  • Do not publish private holdout cases to prompt authors as demonstrations.

Connected lessons

Continue with: Synthetic evaluation holdouts: stop the generator from teaching the test.

prompt engineering
evaluation
Storage details