A random image split can overstate generalization when the same receipt, device or capture session appears on both sides.
Vision evaluation splits: group captures and test acquisition shift
Group near duplicates
A user may upload several crops of one receipt. Those images share text, lighting and device artifacts, so treat the submission as one split unit. Hashes can find exact duplicates, but near-duplicate review still needs a policy. Group validation prevents a model from memorizing the same physical document in train and test.
Add a temporal holdout
Camera software, compression and receipt layouts change. Keep a later capture period untouched while choosing the model, then report performance on it. A model that works on old scanner images may fail on dim mobile photos. Show results by device family, channel, language and document type, with counts and intervals where sample size allows.
Separate label leakage
Filenames, upload paths or processing flags can encode the reviewer’s decision. Strip such metadata from model inputs and verify that a baseline trained on filenames alone cannot predict the label. Apply preprocessing fit only to training data, and use the exact same decoder and resize policy in validation. Augmentation belongs only to training.
Inspect a split fixture
Create three submissions with two images each. All images from one submission must share a partition. Add a later fourth submission captured by a new device and reserve it for a shift test. Report both image count and submission count; a six-image test set with only three independent receipts is not six independent examples.
Implementation
def group_partition(image_records, test_submission_ids):
test_ids = set(test_submission_ids)
training = []
testing = []
for record in image_records:
destination = testing if record["submission_id"] in test_ids else training
destination.append(record["image_id"])
return training, testingPerformance and operating cost
Assigning N images to partitions costs O(N + G) expected time and O(N + G) space for G selected groups. Per-device and temporal slices reduce effective sample sizes; report uncertainty rather than treating every crop as independent evidence.
Common Mistakes
- Do not split near-identical captures across training and evaluation.
- Do not allow label-bearing filenames into model features.
- Do not report only an overall score when acquisition channels differ.
Read next
- Vision dataset contract: image unit, label definition and consent
- Vision augmentation: preserve labels and match serving transforms
- Vision decision metrics: separate localization, class errors and abstention
- Group and time validation: split by the failure you expect in production
Continue the workflow: Negative transfer by target slice.
