Generate text-free receipt-paper textures, then decide whether they add real held-out classifier value without sacrificing privacy or mode coverage.
Project: audit a receipt-paper texture GAN
Limit the generated object
Use paper texture patches with merchant names, dates, amounts, barcodes and account information removed before training. Record the image range, patch size, removal process and physical-receipt split. Keep all crops of each receipt in one split. A tiny GAN can establish the update mechanics on synthetic vectors, but production texture work needs a suitable image generator and sufficient diverse data. Compare with the existing diffusion-background experiment rather than assuming one generative family wins. That project has a matching downstream test.
Train with separate optimizer ownership
Record discriminator and generator revisions, update ratio, noise distribution and fixed-seed visual outputs. During the discriminator update, detach generated patches; during the generator update, retain discriminator input gradients and step only the generator. Save both optimizer states if a run may resume. Monitor real/fake scores and gradient norms but do not treat a familiar score balance as success. The update lesson gives the exact boundary. Reject any patch where text removal has left recognizable private fields.
Audit every mode and near neighbor
Reserve real held-out texture patches for coverage measurement and set a fixed sampling protocol. Compare clean, creased, shadowed and fibrous paper groups, then inspect nearest generated-to-training pairs. Test multiple image feature spaces because a broad feature can miss faint text. Manually review the closest pairs and record exclusions. The code below defines a release gate over measured outcomes; its illustrative numbers are not a claim about a trained generator. Coverage auditing specifies the inputs to that gate.
Run the paired downstream experiment
Train a receipt-quality classifier once on approved real examples alone and once with approved synthetic backgrounds, keeping label budget, model, optimizer and real test identities fixed. Measure blur and clipped-edge recall, false accepts and manual-review load on real held-out captures. Report generated sample counts and how many were rejected in privacy or coverage screening. Synthetic variety does not automatically produce useful variation for a particular classifier; the downstream decision is the final arbiter.
Release a reversible artifact
Package the generator, preprocessing, removal method, seed manifest, nearest-neighbor audit and paired classifier results. Set minimum covered-mode count, maximum unreviewed near-duplicate count and class-specific decision gates before final testing. If any gate fails, keep the generator out of production augmentation. A later generator checkpoint must repeat every gate because adversarial training can gain one mode while losing another. Maintain the previous approved training set and classifier for rollback.
Implementation
from dataclasses import dataclass
@dataclass(frozen=True)
class TextureAudit:
covered_modes: int
private_or_unreviewed_neighbors: int
clipped_edge_recall: float
false_accepts: int
def rejection_reasons(reference: TextureAudit,
candidate: TextureAudit) -> list[str]:
failures = []
if candidate.covered_modes < 4:
failures.append("missing paper mode")
if candidate.private_or_unreviewed_neighbors != 0:
failures.append("unresolved neighbor audit")
if candidate.clipped_edge_recall < reference.clipped_edge_recall:
failures.append("clipped-edge recall")
if candidate.false_accepts > reference.false_accepts:
failures.append("false accepts")
return failures
real_only = TextureAudit(4, 0, 0.93, 2)
synthetic_arm = TextureAudit(3, 0, 0.94, 2)
assert rejection_reasons(real_only, synthetic_arm) == ["missing paper mode"]Performance and operating cost
Training uses two networks and often several discriminator evaluations per generator update; generator inference alone is cheaper than training. The expensive evidence is repeated sampling, nearest-neighbor inspection and paired downstream model training. The gate function is O(1) once its inputs are measured, but calculating exact image-neighbor distances can be O(MRd) for M generated samples, R training images and embedding width d. Report audit sample sizes and device cost instead of presenting illustrative code values as measured results.
Common Mistakes
- Do not train on paper patches that retain private receipt text.
- Do not release after a curated image review without coverage and neighbor counts.
- Do not assume lower generator loss improves real held-out classifier decisions.
