A reproducible analysis identifies the exact input snapshot, transformation version and time boundary that produced a result.
Reproducible analysis snapshots: pin data, code and cutoff together
Freeze the question with the extract
A receipt report generated on Monday may change on Thursday because late reviews arrived or a source table was corrected. A query text alone is not enough to reproduce Monday’s number. Record the extraction cutoff, source snapshot identifier, schema version, transform revision and row counts. If storage supports time travel, use the snapshot reference; otherwise retain a governed immutable extract or a manifest of partition checksums. Metric definitions] belong in the same release record.
Separate rerun from restatement
A rerun against the pinned snapshot should produce the same aggregate. A restatement against newer data is a different result with a reason and a new version. Keep timezone, locale, dependency lockfile and random seed where relevant. A seed does not freeze external data or guarantee bitwise equality across every library and hardware version, so test important invariants rather than promising universal byte identity.
Make the manifest useful
Store counts before and after filtering, key uniqueness checks, missingness summary, and the output artifact digest. Avoid putting credentials or raw personal records in the manifest. Reviewers can then isolate whether a changed result came from data arrivals, code, schema or metric policy. Exploratory checks] should use the same snapshot when they support a published claim.
Verify the rerun
Execute the pipeline twice from the same immutable input, compare aggregate output and record any accepted nondeterminism. Then run it against a deliberately changed partition to ensure the manifest detects a difference. A notebook with hidden state cannot serve as the only execution record; run top to bottom in a clean environment.
Implementation
manifest = {
"snapshot_id": input_snapshot_id,
"cutoff_utc": cutoff.isoformat(),
"transform_revision": build_revision,
"source_rows": len(source_receipts),
"eligible_rows": len(eligible_receipts),
"result_sha256": hashlib.sha256(result_bytes).hexdigest(),
}
assert manifest["eligible_rows"] <= manifest["source_rows"]
manifest_file.write_text(json.dumps(manifest, sort_keys=True, indent=2))Performance and operating cost
Keeping immutable extracts consumes storage proportional to retained data. A small manifest is cheap, while full reruns cost the original query and transformation work; retention should match audit and privacy obligations.
Common Mistakes
- Do not call a changed source table the same snapshot.
- Do not assume a random seed pins data or dependency behavior.
- Do not put secrets or raw personal fields into the manifest.
Read next
- Metric denominators and cohorts: make a rate reproducible
- Exploratory analysis without peeking: inspect the data and preserve the test
- Dataset grain and join cardinality: protect the unit of analysis
- Group and time validation: split by the failure you expect in production
Connected implementation
Continue the workflow: Data Engineering Tutorial.
Continue the workflow: SQL window functions: select one event with a deterministic rule.
Continue the workflow: Instrumentation changes and metric guardrails for product analysis.
Continue the workflow: Measurement resolution and rounding: avoid invented precision.
Continue the workflow: Entity merge lineage: show how identity decisions change metrics.
Continue the workflow: Beta-binomial updating with an auditable case count.
Continue the workflow: Replay manifests and audit trails.
