Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Image augmentation: split originals first and preserve the label

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Augmentation creates training variants only after an identity-aware split, using transforms that keep the label true.

Split originals first

Two photos of one receipt, or crops of one photo, belong in the same partition. Split by receipt identity before generating variants. If augmentation runs first and variants land in training and validation, the model can recognize the underlying receipt instead of generalizing. Grouped validation] addresses the same identity issue for tabular records.

Check label preservation

A mild brightness change may preserve a damage label. A crop removing the torn corner does not. Define allowed transforms per label and inspect a sample of outputs. A transform is not safe merely because it is common in image pipelines. Keep the untouched validation distribution close to deployment input, including blur or poor lighting when those occur.

Make randomness auditable

Record transform names, ranges, library configuration and a run seed. A seed helps repeatability, though worker scheduling and accelerator kernels can still vary. Store original receipt IDs with augmented variants for group checks. Checkpoint metadata] should include the augmentation policy version.

Compare against a baseline

Train once without augmentation and once with the declared policy under the same split and evaluation procedure. Report overall results and the failure slices that motivated the policy. If improvement appears only on augmented validation images, that test cannot support a deployment claim.

Implementation

python
def assert_disjoint_receipts(training_rows, validation_rows):
    training_ids = {row.receipt_id for row in training_rows}
    validation_ids = {row.receipt_id for row in validation_rows}
    overlap = training_ids & validation_ids
    if overlap:
        raise ValueError(f"receipt overlap: {len(overlap)}")

assert_disjoint_receipts(training_rows, validation_rows)
augmented_training_rows = augment_receipts(training_rows, policy_version)

Performance and operating cost

Set overlap checks cost O(N) time and memory for N unique IDs. Offline augmentation adds storage proportional to variants; online augmentation spends CPU per batch and can limit loader throughput.

Common Mistakes

  • Do not split variants independently of their original.
  • Do not crop away evidence while retaining the original label.
  • Do not augment validation to inflate the result.

Read next

Continue the workflow: Pretraining overlap and provenance audit.

ai-data
deep-learning
Storage details