A noise predictor is trained against a precisely defined corruption schedule; changing the schedule or target at inference creates a different problem.
Diffusion forward noise and timestep targets
Define the clean-data contract
Start with a frozen image transform, value range and channel order. For receipt backgrounds, that may mean one grayscale channel scaled to minus one through one, with text and personal information removed before training. Split physical receipts before creating crops. The clean image x0, sampled Gaussian noise epsilon and integer timestep t are three distinct pieces of the training record. Save the transform version beside the noise schedule. A model trained on centered pixels will not interpret raw byte-valued pixels correctly. The receipt classifier uses its own preprocessing contract.
Derive one noisy image directly
Choose beta at each training timestep and define alpha as one minus beta. The cumulative product alpha_bar[t] records how much clean signal remains after t. Then x_t equals square-root alpha_bar[t] times x0 plus square-root one-minus-alpha_bar[t] times epsilon. Sampling x_t in one expression avoids iterating through every earlier corruption step during training. It also makes an exact reconstruction check possible when x0 and epsilon are retained. A different beta table changes the meaning of the same integer timestep, even if image shape and model weights match.
Choose the prediction target explicitly
A basic noise-prediction objective asks the network for epsilon and scores its output against the exact sampled epsilon, often with mean squared error. Other parameterizations predict a clean sample or a velocity-like combination; their conversions and loss weights differ. Store the target type in the artifact and never infer it from a filename. Timestep embeddings must receive the same indexing convention at train and inference time. If one implementation calls the first step zero and another calls it one, their noise levels no longer line up. Loss contracts matter more than a convenient tensor shape.
Inspect the schedule endpoints
Inspect the lowest and highest alpha_bar values as well as images at early, middle and late steps. A late image that still exposes every receipt line may not be a useful near-noise endpoint; an early image that loses text structure immediately may make small-step training harder. Do not judge the schedule only from a graph of beta values. Check that every beta lies strictly between zero and one, that alpha_bar decreases, and that the calculation stays finite in the chosen precision. Keep schedule construction in full precision if a low-precision cumulative product underflows.
Audit the sampling distribution
Uniformly sampled timesteps give every index equal training frequency, but equal frequency is not equal difficulty. Report loss by timestep bucket rather than only one average. Record the seed, sampled timestep, noise tensor identity and image identity for a small reproducible audit batch. Augmentation must happen before corruption so the target noise still matches the image being scored. If a training transform edits x_t after epsilon was drawn, the original epsilon is no longer the exact target. Reverse stepping relies on this same schedule.
Implementation
import torch
torch.manual_seed(47)
clean_receipts = torch.rand(3, 1, 8, 8) * 2 - 1
beta_schedule = torch.linspace(0.001, 0.047, 47, dtype=torch.float64)
alpha_retained = torch.cumprod(1.0 - beta_schedule, dim=0)
sampled_steps = torch.tensor([4, 23, 46])
signal = alpha_retained[sampled_steps].sqrt().view(-1, 1, 1, 1)
noise_scale = (1 - alpha_retained[sampled_steps]).sqrt().view(-1, 1, 1, 1)
sampled_noise = torch.randn_like(clean_receipts)
noisy_receipts = signal * clean_receipts + noise_scale * sampled_noise
recovered_clean = (noisy_receipts - noise_scale * sampled_noise) / signal
torch.testing.assert_close(recovered_clean, clean_receipts, atol=1e-6, rtol=1e-6)
assert bool(torch.all(alpha_retained[1:] < alpha_retained[:-1]))Performance and operating cost
The direct corruption step is O(BCHW) for batch size B and image shape C by H by W, plus O(T) schedule construction for T timesteps. Training cost is dominated by one network forward and backward pass per batch. Keeping a full tensor of independently noised copies for every timestep would multiply image storage by T without improving this objective. The code uses float64 for the schedule audit; production may cast selected coefficients to model dtype after constructing and checking them.
Common Mistakes
- Do not reuse a checkpoint with an altered beta schedule or prediction target.
- Do not normalize x0 twice while leaving epsilon unchanged.
- Do not assume low loss at easy timesteps implies adequate denoising at late steps.
Read next
- Logits, cross-entropy and gradients: align the training calculation
- Training and validation modes: measure the model you will serve
- Diffusion reverse steps and conditional guidance
- Project: train and audit a receipt-background denoiser
- Project: classify receipt image quality with a checked training contract
Continue the workflow: Project: audit a receipt-paper texture GAN.
