Utility is measured by the decision or workflow the generated data supports, using real outcomes that generation never saw.
Synthetic-data utility: downstream tests on an untouched holdout
Define comparison arms
Train or tune the same downstream model on real-only, generated-only and mixed data where policy permits. Keep model architecture, feature pipeline and evaluation horizon fixed. A generated dataset may improve rare-case recall while lowering calibration or performance on ordinary traffic. Report that tradeoff rather than reducing it to one score. Group and time splits protect the real holdout from repeated receipt templates.
Keep the holdout independent
Choose a later time period or unseen source before generator fitting and parameter tuning. Do not use its errors to rewrite the generator repeatedly and still call it untouched. If the holdout is used for iteration, create a new final evaluation set. Record generator version, training-data cutoff and any filters that could move real observations into the synthetic training source.
Check more than marginals
Compare constraints, joint relationships, missingness and task metrics by subgroup. Histograms can agree while a merchant-tax dependency is inverted. For a parser, compare field accuracy and manual-review burden; for forecasting, compare error by horizon and source. Scenario tests identify known failures, but they cannot measure unknown real-world shifts.
Work through a result
Suppose a generated-only parser recovers 42 of 47 fields in a new-scanner holdout while a real-only baseline recovers 44. The generated set still may be useful if it catches a refund edge case, but do not claim general improvement. Inspect errors, confidence and support counts. Keep a predeclared release threshold so a selectively reported slice cannot override the full result.
Implementation
def accuracy_by_source(predictions):
totals = {}
correct = {}
for result in predictions:
source = result["source"]
totals[source] = totals.get(source, 0) + 1
correct[source] = correct.get(source, 0) + int(result["predicted"] == result["actual"])
return {source: correct[source] / count for source, count in totals.items()}Performance and operating cost
Grouping N predictions costs O(N) time and O(S) space for S sources. Training and evaluation costs depend on the downstream model; repeated holdout use adds statistical selection cost even when runtime is cheap.
Common Mistakes
- Do not evaluate on records used to fit or tune the generator.
- Do not substitute visual similarity for downstream utility.
- Do not hide subgroup regressions behind a pooled metric.
