Batch construction controls memory, gradient noise and class exposure; sampling changes the training distribution.
Batching and class sampling: know the population the optimizer sees
Choose size by measurement
A batch must fit inputs, activations, optimizer state and temporary workspace. Start with a measured size and observe peak memory and examples per second. Doubling the batch can improve device use, but can also reduce update frequency or exhaust memory. Keep the short final batch or say why it is discarded; repeated dropping can exclude particular records when order is stable.
Treat imbalance as policy
Oversampling rare damaged receipts makes the optimizer see a class mix unlike live traffic. That may improve ranking of unusual cases, but raw probabilities no longer automatically describe deployment prevalence. Evaluate on a holdout with the actual mix and calibrate if decisions require probabilities. Thresholds] should reflect review capacity and error costs.
Keep identities apart
A sampler can put near-duplicate photos of one receipt into many training batches. That is acceptable within training, but not across training and validation. Split by receipt or customer group before creating samplers or variants. The split boundary] defines the correct order.
Audit effective exposure
Count unique receipt IDs and class counts over a sampled epoch. Check every target index and confirm that no evaluation ID appears in training. Document replacement. When a rare class has only a few unique cases, repeated exposure is not new information; report the unique count beside sample count.
Implementation
from torch.utils.data import WeightedRandomSampler
training_counts = torch.bincount(training_class_ids, minlength=class_count)
if (training_counts == 0).any():
raise ValueError("training class has no observations")
per_receipt_weight = 1.0 / training_counts[training_class_ids].float()
training_sampler = WeightedRandomSampler(
weights=per_receipt_weight,
num_samples=len(training_class_ids),
replacement=True,
)Performance and operating cost
A batch of B images adds O(BCHW) input memory plus model activations. Constructing weights is O(N); sampling N records per epoch adds selection work and can repeat scarce rows many times.
Common Mistakes
- Do not assume oversampled probabilities match live prevalence.
- Do not sample before a grouped split.
- Do not discard a short batch silently.
