Build a receipt-quality classifier with explicit effective-batch accounting, a measured precision decision and recovery evidence rather than hiding memory failures behind a smaller batch.
Project: train a receipt model under a memory budget
Fix the experimental contract
Split receipt identities before augmentation and freeze a label map for readable, blurred and cut-off images. Record the device memory limit, image dimensions, model revision, optimizer, loss reduction and number of examples per optimizer update. Keep one full-precision baseline small enough to run. The original receipt project supplies the data and evaluation boundary; this project tests whether the optimized training loop preserves it.
Build a weighted accumulation loop
Choose a microbatch that fits the device, then accumulate to an effective batch of 47 accepted examples. The final microbatch will often be short; sum unreduced losses and divide by the actual number of accepted examples, not by a rounded-up microbatch count. Clip only after all contributions are present, and step once. Test a copied model on a tiny deterministic batch against a full-batch gradient before using the full image dataset. Effective-batch accounting is a release check.
Introduce precision after the baseline
Measure full precision, float16 autocast with scaling and, if supported, bfloat16 autocast. Record peak allocated memory, accepted examples per second, nonfinite gradients, skipped steps and frozen validation metrics. If precision changes predictions near the acceptance threshold, quantify the change by class. Scaling and clipping must occur in the documented order.
Interrupt and restore
Save a checkpoint only at an optimizer-step boundary with model, optimizer, scheduler, scaler, sampler position and step count. Crash after an unfinished checkpoint write and prove the last complete checkpoint still resumes. Compare the first resumed fixed-batch update against an uninterrupted reference under a controlled seed. A new device type or changed microbatch plan may prevent bitwise identity; state the expected tolerance.
Report the tradeoff
Submit a memory trace, training-time breakdown, gradient comparison, update and skip counts, validation result and checkpoint recovery trace. Also show a deliberately wrong run that divides a short microbatch mean by the number of microbatches; report its gradient discrepancy. Do not declare success solely because an out-of-memory error disappeared while the accepted-example count or class mix changed.
Implementation
import copy
import torch
from torch import nn
torch.manual_seed(83)
reference_model = nn.Linear(3, 3)
bounded_model = copy.deepcopy(reference_model)
receipt_vectors = torch.tensor([[0.8, 0.1, 0.1],
[0.2, 0.5, 0.3],
[0.1, 0.3, 0.6]])
quality_targets = torch.tensor([0, 1, 2])
nn.functional.cross_entropy(reference_model(receipt_vectors),
quality_targets).backward()
for start, stop in ((0, 2), (2, 3)):
loss_sum = nn.functional.cross_entropy(
bounded_model(receipt_vectors[start:stop]),
quality_targets[start:stop], reduction="sum")
(loss_sum / quality_targets.numel()).backward()
for reference_weight, bounded_weight in zip(reference_model.parameters(),
bounded_model.parameters()):
torch.testing.assert_close(reference_weight.grad, bounded_weight.grad,
rtol=1e-6, atol=1e-7)Performance and operating cost
Peak activation memory follows the largest microbatch, but model, gradients and optimizer moments still occupy memory. More microbatches can lower examples per second through repeated launches, while autocast may restore throughput on suitable hardware. Instrument both peak memory and wall-clock cost. A checkpoint stores additional optimizer and scaler state, and final selection must use the same validation population as the baseline.
Common Mistakes
- Do not change the dataset split while comparing precision modes.
- Do not omit the short final microbatch or give it the same weight as a full microbatch.
- Do not resume from weights alone and call the optimizer trajectory recovered.
Read next
- Gradient accumulation and effective batch accounting
- Mixed precision, loss scaling and gradient clipping
- Project: classify receipt image quality with a checked training contract
- Checkpoint recovery: save optimizer state and the run boundary
- Image augmentation: split originals first and preserve the label
Continue the workflow: Project: localize receipt defects with a residual CNN.
Continue the workflow: Project: prove multiworker receipt-training parity.
Continue the workflow: Project: select a receipt training schedule without test leakage.
Continue the workflow: Project: measure activation recomputation on receipt training.
