A reference set should contain reviewed decisions with reasons and history, not a mutable spreadsheet of final answers.
Adjudication and gold sets: versioned reference decisions
Define who decides
When two reviewers disagree, a third reviewer can inspect both reasons while making a documented final decision. Escalate policy gaps to a guideline owner rather than forcing a binary choice. Record whether the decision is evidence-based, policy-based or unresolved. A majority vote is not sufficient when all reviewers saw the same misleading crop or misunderstood the same rule.
Freeze evaluation versions
Publish a gold-set version with item IDs, evidence revisions, guideline version, adjudication outcomes and exclusions. A model evaluation must name that version. If an item is corrected later, create a new version and report the delta; changing labels in place makes historic model comparisons impossible. Validation boundaries still apply: near-duplicate receipts cannot be split across training and a final test.
Review disagreement clusters
Sample adjudicated items by scanner source and label class. A reviewer may be consistently wrong for a specific receipt layout, while the aggregate agreement looks fine. Keep the original reviews visible to authorized auditors. Repeated disagreements of the same type signal that the guideline needs a sharper boundary or the interface needs more evidence, not that more voting is automatically useful.
Run a correction drill
Freeze 47 items as gold version 3. A later scan reveals two totals were cropped and one class rule was applied incorrectly. Publish version 4 with those three item-level corrections and reasons, then rerun a held-out model against both versions. The reported score difference should identify label corrections separately from a model change.
Implementation
def gold_delta(previous, current):
return {receipt_id: (previous.get(receipt_id), current.get(receipt_id))
for receipt_id in previous.keys() | current.keys()
if previous.get(receipt_id) != current.get(receipt_id)}Performance and operating cost
Comparing G item IDs costs O(G) expected time and O(D) output space for D changes. Versioned evidence and decisions consume more storage, but make test-score changes and label corrections auditable.
Common Mistakes
- Do not overwrite a frozen gold set in place.
- Do not resolve a policy ambiguity by unexplained majority vote.
- Do not mix near duplicates between model training and final evaluation.
