Build a small conditional noise predictor for synthetic receipt backgrounds, then reject memorization and prove downstream value on real held-out captures.
Project: train and audit a receipt-background denoiser
Specify the limited target
Produce grayscale background texture only, with text, merchant names, dates and identifiers removed before any model training. Hold out entire physical receipts, not crops, and freeze a 47-step schedule with the exact input scaling and condition vocabulary. The condition in this exercise is a coarse paper-contrast value. The 8 by 8 network below demonstrates one noise-prediction update, not a production image generator; real texture synthesis needs more receptive field, data and evaluation. The forward-noise contract establishes the target.
Train both condition modes
Draw a clean synthetic background, sample an integer timestep and fresh Gaussian noise, then construct x_t. Append the normalized timestep and contrast condition to the flattened noisy image. Drop the contrast condition for a recorded subset of updates and supply a separate null flag, so the network can distinguish missing conditioning from a genuine zero-contrast request. Score predicted epsilon against the exact sampled epsilon. Log loss by timestep bucket and condition-present status; an aggregate loss can hide a branch that never learns.
Build a fixed-noise sampling audit
After training, sample from fixed seeded noise using one documented reverse trajectory. Compare guidance scale one against a stronger scale while holding the noise seed, model and schedule constant. Inspect range violations, near-constant outputs and diversity across seeds. Do not advertise low training loss as proof of image quality. Record model revision, preprocessing and scheduler config with every generated set. Reverse guidance describes the two predictions required at each step.
Check privacy and duplication
Search generated backgrounds for nearest training-image neighbors using several measures, including raw-pixel distance after alignment and a feature representation. Manually review the closest pairs; a distance threshold alone does not certify absence of memorized text. Keep an immutable list of excluded personal fields and check source permissions before training. If background removal left a faint account number, synthetic release must stop even if the denoising objective is numerically correct. Distinct synthetic seeds are not proof of distinct real-world information.
Measure a downstream decision
Train the same receipt-quality classifier twice: once with real data alone, once with the approved synthetic backgrounds used only as training augmentation. Keep the real held-out receipt set and threshold-selection policy fixed. Report blur and cut-off recall, false accepts, calibration and device slices. Reject the synthetic set if it worsens an agreed safety slice or offers no repeatable gain. Deliver the schedule, training manifest, nearest-neighbor audit and paired classifier results. Confidence review can expose harmful shifts masked by one accuracy figure.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
backgrounds = torch.rand(4, 1, 8, 8) * 2 - 1
paper_contrast = torch.tensor([[0.2], [0.7], [0.4], [0.9]])
beta_schedule = torch.linspace(0.001, 0.047, 47)
alpha_retained = torch.cumprod(1 - beta_schedule, dim=0)
sampled_steps = torch.randint(0, 47, (4,))
sampled_noise = torch.randn_like(backgrounds)
signal = alpha_retained[sampled_steps].sqrt().view(4, 1, 1, 1)
noise_scale = (1 - alpha_retained[sampled_steps]).sqrt().view(4, 1, 1, 1)
noisy_backgrounds = signal * backgrounds + noise_scale * sampled_noise
condition_present = torch.tensor([[1.0], [0.0], [1.0], [0.0]])
condition_value = paper_contrast * condition_present
model_inputs = torch.cat((noisy_backgrounds.flatten(1),
sampled_steps.float().unsqueeze(1) / 46,
condition_value, condition_present), dim=1)
noise_predictor = nn.Sequential(nn.Linear(67, 96), nn.ReLU(), nn.Linear(96, 64))
optimizer = torch.optim.AdamW(noise_predictor.parameters(), lr=0.001)
optimizer.zero_grad(set_to_none=True)
predicted_noise = noise_predictor(model_inputs).view_as(sampled_noise)
training_loss = functional.mse_loss(predicted_noise, sampled_noise)
training_loss.backward()
optimizer.step()
assert torch.isfinite(training_loss)Performance and operating cost
This illustrative dense predictor has forward cost proportional to the input width times 96 plus 96 times 64 per sample, and it stores activations for backpropagation. A realistic image denoiser is much more expensive and repeats a forward pass for every sampling step; conditional guidance typically adds another pass. Nearest-neighbor checks also require a search over generated and training images, so define an audit sample and indexing plan before scaling output volume. The project accepts no speed or privacy claim from the tiny code example.
Common Mistakes
- Do not train on receipt text and assume later generation is private.
- Do not reuse validation receipts as generator training data or augmentation seeds.
- Do not infer downstream utility from denoising loss or visual appeal alone.
Read next
- Diffusion forward noise and timestep targets
- Diffusion reverse steps and conditional guidance
- Project: classify receipt image quality with a checked training contract
- Project: audit receipt confidence and review handoff
- Staged unfreezing and domain-shift audits
Continue the workflow: Project: audit a receipt-paper texture GAN.
