Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Diffusion reverse steps and conditional guidance

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A reverse update combines the predicted noise with the exact training schedule; guidance adds a second prediction whose training meaning must be known.

Recover a clean estimate first

For an epsilon-predicting model, estimate x0 by subtracting the predicted noise contribution from x_t and dividing by the retained-signal coefficient. This estimate is not guaranteed to lie in the training image range, particularly under strong guidance or at a high-noise step. Decide whether and where to clip it, then keep that choice fixed during evaluation. A deterministic reverse update can combine the clean estimate and predicted noise at an earlier alpha_bar value. It is one sampler rule, not a claim that all diffusion samplers are interchangeable. The forward schedule defines these coefficients.

Define conditional and unconditional branches

Classifier-free guidance requires a predictor that has learned to operate both with and without the condition. During training, deliberately replace a documented fraction of conditions with a null condition; during sampling, run the same model with the null and requested conditions. Combine outputs as unconditional plus guidance-scale times conditional-minus-unconditional. A scale of one yields the conditional prediction in this convention. An arbitrary empty string or zero vector is not automatically the trained null representation. For structured receipt metadata, reserve an explicit null token and test it.

Price the guidance choice

The two branches can often be evaluated in one doubled batch, but they still perform roughly twice the model work and may raise peak memory. Raising guidance can make requested attributes more visible while reducing variety or pushing values outside the trained range. Evaluate fidelity and diversity together. If the condition names a receipt defect, inspect whether generated images contain that defect rather than merely shifting brightness. A guidance scale tuned for one schedule, resolution or model revision is not a universal number.

Keep the timestep trajectory explicit

A sampler may skip training indices to reduce latency. Store the ordered inference indices, the scheduler rule and endpoint convention with the generated artifact. A jump from step 46 to 23 is not the same update as 46 to 45. The code below demonstrates a deterministic transition from step 23 to 22 and does not add stochastic posterior variance. It uses dummy predictions to check algebra; a trained network must supply them in deployment. Run a fixed-noise replay after any scheduler change and compare output distributions, not just one pleasing image.

Gate synthetic data use

Generated receipt backgrounds can help stress a downstream quality classifier, but they must not leak holdout receipts or replace evaluation on real captures. Audit near-duplicates against training inputs, visible personal fields, class labels and capture-device mix. Report downstream error on real held-out receipts with and without generated augmentation. If synthetic training improves average accuracy while making a rare cut-off class worse, the project fails its intended purpose. The applied project makes these gates concrete.

Implementation

python
import torch

beta_schedule = torch.linspace(0.001, 0.047, 47, dtype=torch.float64)
alpha_retained = torch.cumprod(1 - beta_schedule, dim=0)
noisy_receipt = torch.full((1, 1, 8, 8), 0.18, dtype=torch.float64)
unconditional_noise = torch.full_like(noisy_receipt, 0.04)
conditional_noise = torch.full_like(noisy_receipt, 0.07)
guidance_scale = 1.7
guided_noise = unconditional_noise + guidance_scale * (
    conditional_noise - unconditional_noise)
current_step, previous_step = 23, 22
clean_estimate = (noisy_receipt -
                  (1 - alpha_retained[current_step]).sqrt() * guided_noise
                  ) / alpha_retained[current_step].sqrt()
previous_receipt = (alpha_retained[previous_step].sqrt() * clean_estimate +
                    (1 - alpha_retained[previous_step]).sqrt() * guided_noise)
assert previous_receipt.shape == noisy_receipt.shape
assert torch.isfinite(previous_receipt).all()

Performance and operating cost

For K reverse steps, one conditional pass per step costs roughly K model forwards; guidance normally needs two predictions per step, so model compute is near 2K forwards. Activation batching may improve throughput without eliminating the work. Memory depends on implementation: a combined conditional/unconditional batch can double activations, while sequential calls add latency. The algebraic tensor operations shown here are O(BCHW) per step and are usually much cheaper than the denoiser.

Common Mistakes

  • Do not request an unconditional prediction from a model never trained with null conditions.
  • Do not swap timestep schedules while keeping the old model checkpoint.
  • Do not treat a deterministic one-step example as a complete stochastic sampler.

Read next

ai-data
deep-learning
Storage details