Classification training needs raw class scores, integer targets and an explicit gradient lifecycle.
Logits, cross-entropy and gradients: align the training calculation
Follow the score contract
A classifier produces one raw score per class. Cross-entropy consumes those logits and class-index targets; it performs the stable normalization needed for the loss. Applying softmax before that loss compresses scores a second time and changes gradients. Convert logits to probabilities only when a metric or decision policy needs them. Calibration] tests whether those probabilities match observed frequency.
Control one update
For each batch, clear old gradients, run the forward pass, calculate loss, backpropagate and step the optimizer. If gradients are deliberately accumulated, divide loss or adjust the update policy and record how many examples contribute. An omitted zero step otherwise accumulates gradients by accident. Variable batch sizes can also change the effective weight of an update.
Inspect the failure path
Verify finite loss and gradient norms, target range and output shape. A falling training loss is insufficient evidence of generalization: the model may learn a leaked field or repeated receipt identity. Grouped and time-aware validation] makes that shortcut more visible. Stop rather than silently skipping every non-finite batch.
Overfit a small batch first
Train on a fixed tiny batch until the model can fit it. If it cannot, inspect labels, learning rate, loss inputs and gradient flow before growing the architecture. Then test a second fixed batch size. This narrow experiment catches shape and optimizer mistakes faster than a full training run.
Implementation
optimizer.zero_grad(set_to_none=True)
logits = receipt_model(image_batch)
if logits.shape != (image_batch.shape[0], class_count):
raise ValueError("class-score shape changed")
loss = torch.nn.functional.cross_entropy(logits, class_ids)
if not torch.isfinite(loss):
raise ValueError("non-finite loss")
loss.backward()
optimizer.step()Performance and operating cost
Forward and backward passes dominate computation and activation memory; backward keeps intermediates related to batch size and model depth. A larger batch helps throughput only until memory or generalization becomes the limit.
Common Mistakes
- Do not softmax before a cross-entropy API expecting logits.
- Do not accumulate gradients accidentally.
- Do not treat lower training loss as proof of deployment value.
Read next
- Tensor contracts: shape, dtype, device and mask
- Training and validation modes: measure the model you will serve
- Checkpoint recovery: save optimizer state and the run boundary
- Probability calibration: test whether risk scores mean what they say
Continue the workflow: Gradient accumulation and effective batch accounting.
Continue the workflow: Project: localize receipt defects with a residual CNN.
Continue the workflow: Variational latent sampling, KL accounting and collapse.
Continue the workflow: Temperature scaling and selective risk.
Continue the workflow: Diffusion forward noise and timestep targets.
Continue the workflow: Distillation soft targets, temperature and label balance.
Continue the workflow: GAN discriminator and generator update boundaries.
