Controlled synthetic scenarios expose specific failure modes, but their frequency and appearance need not match live traffic.
Rare-case simulation: scenario coverage and domain gap
Build a scenario matrix
List combinations that matter to the consumer: cropped totals, duplicate uploads, mixed tax rates, delayed refunds and scanner-specific blur. Mark expected behavior for each, including rejection and manual review. Pair ordinary cases with extreme values so a parser is not tuned only to the spectacular failures. Label rules should classify every generated scenario without using its generator tag as a shortcut.
Vary causes, not just values
Changing a price from 47 to 48 tests one numeric range. Changing whether a discount applies before tax tests logic. Model source-specific missingness, order changes and retries when those mechanisms cause real failures. Hold some generator templates out of development tests; otherwise a model can memorize the synthetic style and appear to generalize while failing on a new scanner.
Quantify coverage honestly
Report how many scenarios exist, which constraints they hit and which consumers passed them. Distinguish scenario coverage from population prevalence. A generated set with 30% refunds may be excellent for exercising refund logic yet useless for estimating the true refund rate. Real holdout evaluation is needed before making a claim about deployment performance.
Run an adversarial batch
Generate 47 receipts: 20 ordinary, nine discounted, eight scanned twice, six with delayed refunds and four with cropped totals. Test a pipeline for duplicate handling, reconciliation and manual-review routing. Then replace the text template and scanner dimensions while keeping the same semantic scenarios. A large score drop reveals template dependence rather than a new business rule.
Implementation
def scenario_counts(generated_rows):
counts = {}
for receipt in generated_rows:
scenario = receipt["scenario"]
counts[scenario] = counts.get(scenario, 0) + 1
return countsPerformance and operating cost
Counting N generated rows costs O(N) time and O(S) space for S scenarios. Exhaustive combinations grow multiplicatively with independent factors, so prioritize by risk and known incident patterns.
Common Mistakes
- Do not treat scenario proportions as real prevalence.
- Do not test only against the same templates used during development.
- Do not use a generator tag as an input feature for the model being evaluated.
