Augmentation creates training variants only after an identity-aware split, using transforms that keep the label true.
Image augmentation: split originals first and preserve the label
Split originals first
Two photos of one receipt, or crops of one photo, belong in the same partition. Split by receipt identity before generating variants. If augmentation runs first and variants land in training and validation, the model can recognize the underlying receipt instead of generalizing. Grouped validation] addresses the same identity issue for tabular records.
Check label preservation
A mild brightness change may preserve a damage label. A crop removing the torn corner does not. Define allowed transforms per label and inspect a sample of outputs. A transform is not safe merely because it is common in image pipelines. Keep the untouched validation distribution close to deployment input, including blur or poor lighting when those occur.
Make randomness auditable
Record transform names, ranges, library configuration and a run seed. A seed helps repeatability, though worker scheduling and accelerator kernels can still vary. Store original receipt IDs with augmented variants for group checks. Checkpoint metadata] should include the augmentation policy version.
Compare against a baseline
Train once without augmentation and once with the declared policy under the same split and evaluation procedure. Report overall results and the failure slices that motivated the policy. If improvement appears only on augmented validation images, that test cannot support a deployment claim.
Implementation
def assert_disjoint_receipts(training_rows, validation_rows):
training_ids = {row.receipt_id for row in training_rows}
validation_ids = {row.receipt_id for row in validation_rows}
overlap = training_ids & validation_ids
if overlap:
raise ValueError(f"receipt overlap: {len(overlap)}")
assert_disjoint_receipts(training_rows, validation_rows)
augmented_training_rows = augment_receipts(training_rows, policy_version)Performance and operating cost
Set overlap checks cost O(N) time and memory for N unique IDs. Offline augmentation adds storage proportional to variants; online augmentation spends CPU per batch and can limit loader throughput.
Common Mistakes
- Do not split variants independently of their original.
- Do not crop away evidence while retaining the original label.
- Do not augment validation to inflate the result.
Read next
- Batching and class sampling: know the population the optimizer sees
- Training and validation modes: measure the model you will serve
- Group and time validation: split by the failure you expect in production
- Project: classify receipt image quality with a checked training contract
Continue the workflow: Pretraining overlap and provenance audit.
