A replay manifest pins the input positions, dataset versions, code version, interval and output identity needed to rerun or explain a pipeline result.
Replay manifests and audit trails
Capture the actual inputs
A day-47 billing run reads source snapshot 41, exchange-rate snapshot 12 and source positions 900 through 1,147. Record those identifiers at execution, not just the logical table names. If the upstream table changes next week, 'rerun day 47' without a pinned version does not reproduce the original result. Snapshot identity anchors file-based tables.
Record transformation identity
Store the code commit or artifact digest, configuration version, schema versions and execution parameters. Secrets do not belong in the manifest. A retry with the same manifest should produce the same canonical output checksum; a deliberate correction should create a new manifest and link to the superseded run.
Represent counts and decisions
For each stage, record extracted, accepted, quarantined, deleted and published counts at a declared grain. A source count of 1,147 with 1,139 accepted, six quarantined and two deletion events can reconcile if those categories partition the source events; explain that arithmetic in the contract. Reconciliation should fail before publication when counts do not balance.
Tie manifests to lineage
The manifest records a run's exact inputs and output. The lineage graph connects that output to later consumers. Together they answer which reports used a bad exchange-rate version and which input should be replayed. Impact analysis walks this graph, then the manifests select the precise intervals.
Protect the audit record
Write manifests immutably or version them with a tamper-evident digest, restrict who can amend them, and retain them at least as long as their associated data and audit requirement. A manifest that names expired source data cannot guarantee replay; keep retention schedules aligned. Document a forward repair when exact replay is impossible, rather than calling a reconstruction identical.
Implementation
import hashlib
import json
run_manifest = {
"interval": [47, 48], "source_snapshot": "snap-41",
"rate_snapshot": "rates-12", "source_positions": [900, 1147],
"code_digest": "billing-v3", "output_snapshot": "snap-42",
}
def manifest_digest(manifest):
canonical = json.dumps(manifest, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(canonical.encode()).hexdigest()
assert manifest_digest(run_manifest) == manifest_digest(dict(reversed(list(run_manifest.items()))))Performance and operating cost
Canonical serialization and hashing cost O(B) time and space for B manifest bytes. The larger operational cost is retaining addressable input versions and connecting manifests to published outputs. A digest detects change; it does not prove the source data was truthful.
Common Mistakes
- Do not store only table names when exact snapshots or offsets determine the result.
- Do not put credentials or raw sensitive payloads in a manifest.
- Do not promise replay after the source retention horizon has passed.
Read next
- DAG intervals and idempotent task outputs
- Pipeline lineage and impact analysis
- Pipeline SLOs, freshness and error budgets
- Backfills: rebuild history without exposing a half-written result
- Project: operate a daily billing pipeline
- Reproducible analysis snapshots: pin data, code and cutoff together
Continue the workflow: Project: release a reconciled revenue mart.
