Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Replay manifests and audit trails

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A replay manifest pins the input positions, dataset versions, code version, interval and output identity needed to rerun or explain a pipeline result.

Capture the actual inputs

A day-47 billing run reads source snapshot 41, exchange-rate snapshot 12 and source positions 900 through 1,147. Record those identifiers at execution, not just the logical table names. If the upstream table changes next week, 'rerun day 47' without a pinned version does not reproduce the original result. Snapshot identity anchors file-based tables.

Record transformation identity

Store the code commit or artifact digest, configuration version, schema versions and execution parameters. Secrets do not belong in the manifest. A retry with the same manifest should produce the same canonical output checksum; a deliberate correction should create a new manifest and link to the superseded run.

Represent counts and decisions

For each stage, record extracted, accepted, quarantined, deleted and published counts at a declared grain. A source count of 1,147 with 1,139 accepted, six quarantined and two deletion events can reconcile if those categories partition the source events; explain that arithmetic in the contract. Reconciliation should fail before publication when counts do not balance.

Tie manifests to lineage

The manifest records a run's exact inputs and output. The lineage graph connects that output to later consumers. Together they answer which reports used a bad exchange-rate version and which input should be replayed. Impact analysis walks this graph, then the manifests select the precise intervals.

Protect the audit record

Write manifests immutably or version them with a tamper-evident digest, restrict who can amend them, and retain them at least as long as their associated data and audit requirement. A manifest that names expired source data cannot guarantee replay; keep retention schedules aligned. Document a forward repair when exact replay is impossible, rather than calling a reconstruction identical.

Implementation

python
import hashlib
import json

run_manifest = {
    "interval": [47, 48], "source_snapshot": "snap-41",
    "rate_snapshot": "rates-12", "source_positions": [900, 1147],
    "code_digest": "billing-v3", "output_snapshot": "snap-42",
}

def manifest_digest(manifest):
    canonical = json.dumps(manifest, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(canonical.encode()).hexdigest()

assert manifest_digest(run_manifest) == manifest_digest(dict(reversed(list(run_manifest.items()))))

Performance and operating cost

Canonical serialization and hashing cost O(B) time and space for B manifest bytes. The larger operational cost is retaining addressable input versions and connecting manifests to published outputs. A digest detects change; it does not prove the source data was truthful.

Common Mistakes

  • Do not store only table names when exact snapshots or offsets determine the result.
  • Do not put credentials or raw sensitive payloads in a manifest.
  • Do not promise replay after the source retention horizon has passed.

Read next

Continue the workflow: Project: release a reconciled revenue mart.

ai-data
data-engineering
Storage details