A label is a recorded decision about a specified unit, at a specified time, under a versioned rule set.
Label contracts: decision unit, taxonomy and review rules
Fix the decision unit
For receipt extraction, decide whether a reviewer labels an entire receipt, one line item or one field. A field-level “merchant missing” label cannot be compared with a receipt-level “unreadable” label without an explicit mapping. Record the document version, crop or field span, event time and reviewer view. Geometry quality belongs in this contract when images are involved.
Write a usable taxonomy
Separate positive, negative, unknown and not-applicable states. Unknown means the evidence cannot settle the question; it is not a negative outcome. Give each class inclusion rules, exclusions and a counterexample. A policy that says “mark suspicious receipts” leaves reviewers to invent different thresholds. Tie every label to a guideline revision so the same fixture can be replayed under a later policy.
Control what reviewers see
Hide downstream model scores and final business outcomes when those facts would bias a primary annotation. If a reviewer must see a model suggestion, store that as assisted review and measure it separately. Mask unnecessary personal data. Data minimization makes the review process safer and keeps the decision tied to evidence actually available at labeling time.
Try boundary cases
Create one clean receipt, one cropped total, one duplicate scan and one receipt whose total is ambiguous because a discount line is cut off. Require a class and reason for each. Run two guideline versions and list precisely which labels change. If a change cannot be explained by a rule, the taxonomy is incomplete or the evidence presented to reviewers differs.
Implementation
LABELS = {"accepted", "rejected", "unknown", "not_applicable"}
def valid_review(review):
return (review["label"] in LABELS
and bool(review["guideline_version"])
and bool(review["receipt_id"])
and bool(review["evidence_revision"]))Performance and operating cost
Contract validation is O(1) per review and O(N) for N reviews. Keeping evidence and guideline revisions increases storage, but prevents silent reinterpretation of historical labels when a policy changes.
Common Mistakes
- Do not merge unknown into the negative class.
- Do not change a guideline without recording its version.
- Do not label a crop as if the reviewer saw the whole document.
