Skip to content
AITroveRead. Build. Understand.
Make this comfortable

GAN mode coverage and near-duplicate audits

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A generator may fool its current discriminator while omitting entire parts of the target population or reproducing training images too closely.

Separate fidelity from coverage

A small number of attractive samples says little about how much of the real population is represented. Partition a held-out real set into meaningful modes such as clean paper, folded edge, low exposure and textured background. Compare generated counts and feature distributions for every mode. These categories need not be training labels; they are an audit vocabulary for missing regions. A generator that emits only clean paper can look polished and still be useless for stress-testing a receipt classifier. The training loss does not expose this omission.

Fix sampling before comparison

Use a fixed seed list and a separately sampled seed list. The fixed set shows whether a checkpoint changed particular outputs; the fresh set estimates population variety. Keep generator revision, noise distribution, output normalization and sample count in the audit record. Never select the most flattering samples by hand for the coverage calculation. Report both within-set nearest-neighbor distances and distances to real training images. Low within-set distance suggests repetition; suspiciously low training-neighbor distance can signal memorization, though distance alone is not a privacy guarantee.

Choose features that match the task

Pixel distance can be dominated by small alignment changes, while an image embedding may ignore a narrow cut-off edge. Use several comparisons: aligned raw pixels, a frozen general image feature and a receipt-quality feature. Inspect nearest pairs manually, particularly any generated sample resembling a training receipt with residual personal text. Do not train the audit feature on the final held-out set. The code below uses two-dimensional audit features to make assignments visible; production images need a validated feature extractor and a human-reviewed closest-pair sample.

Detect memorization and mode collapse separately

A generator can collapse to a few novel outputs without copying training images, or copy many training images while covering apparent modes. Count unique generated neighborhoods and covered held-out modes, then inspect closest training neighbors. Compare the same measures at several checkpoints instead of assuming late training is always better. A discriminator overfitted to its own training set may reward a generator that exploits its blind spots. The diffusion project uses a similar privacy and downstream-utility gate despite a different objective.

Use downstream utility as a final gate

For augmentation, compare a receipt-quality model trained on real data alone with one trained on approved synthetic textures, holding real training identities and test data fixed. Report rare-defect recall, false accepts and device slices. Coverage of visual modes is necessary but not sufficient: a generated texture may have unrealistic correlations with defects. If the synthetic arm harms the held-out decision, do not release the samples merely because a feature-space score improved. The applied project packages these audits.

Implementation

python
import torch

heldout_mode_centers = torch.tensor([[0.0, 0.0], [2.7, 0.3], [-1.6, 2.1]])
generated_features = torch.tensor([[0.1, -0.1], [0.2, 0.0],
                                   [2.6, 0.4], [2.9, 0.1],
                                   [0.0, 0.2], [2.8, 0.2]])
training_features = torch.tensor([[0.0, 0.1], [2.4, 0.3], [-1.5, 2.0]])
assigned_modes = torch.cdist(generated_features, heldout_mode_centers).argmin(dim=1)
covered_mode_count = int(torch.unique(assigned_modes).numel())
closest_training_distance = torch.cdist(
    generated_features, training_features).amin(dim=1)
pairwise = torch.cdist(generated_features, generated_features)
pairwise.fill_diagonal_(float("inf"))
closest_generated_distance = pairwise.amin(dim=1)
assert covered_mode_count == 2
assert closest_training_distance.shape == closest_generated_distance.shape == (6,)

Performance and operating cost

For M generated samples, R real audit samples and feature width d, a brute-force real-neighbor scan costs O(MRd) time; a generated-pair scan costs O(M²d) time and O(M²) memory if the full matrix is kept. Chunked distance computation lowers peak memory, while approximate indexes trade exact neighbors for speed. Image encoding may dominate these comparisons. Audit cost grows with checkpoints and seeds, so choose a fixed protocol rather than silently reducing the sample size for weak candidates.

Common Mistakes

  • Do not present a small curated grid as evidence of population coverage.
  • Do not treat a single feature distance threshold as proof of privacy.
  • Do not use the co-trained discriminator as the sole quality judge.

Read next

ai-data
deep-learning
Storage details