Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Run-scoped lineage and release evidence

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Run-scoped lineage records the exact input and output generations used by one execution, rather than only a planned table dependency.

Distinguish design from execution

A planned DAG says that the revenue job reads orders and rates. It cannot prove which order snapshot or rate revision produced a specific published total. Record a run ID, code release, named input generations, output generation and final status for every publish attempt. Static lineage still helps estimate impact, but runtime evidence answers what actually happened.

Capture the committed boundary

An attempt may read data, write candidate files and then fail. Its lineage is useful for debugging, yet it must not be labelled as the lineage of a published dataset. Attach the output generation only after the atomic commit succeeds. Keep failed attempts under their own run IDs. A consumer should be able to follow the current table pointer to one successful run and then to its exact inputs.

Avoid false many-to-many edges

A job that reads customer and invoice tables may write separate customer and invoice summaries. Treating every input as the parent of every output invents dependencies. Capture edges at the output or transform level, including filtered partitions when practical. If the implementation cannot prove a column-level edge, label it as unknown rather than presenting an inferred edge as certain.

Preserve identity across retries

A retry of the same interval is a new attempt under a stable logical run key. Pin source generations or record that inputs changed before retry; otherwise two attempts with the same scheduled time can calculate different answers. The replay manifest should retain both attempts and name the winning publication. This also makes a rollback reviewable.

Query backward from an answer

Given a suspicious dashboard total, start with its serving generation, find the table snapshot and successful run, then inspect input versions and source positions. Test the chain with a deliberately failed attempt between two successful releases: the failed candidate must appear in diagnostics but never in the published answer. Keep access controls on lineage metadata because identifiers and partition values can reveal sensitive activity.

Implementation

python
runs = [
    {"id": "run-71", "status": "failed", "inputs": ("orders-47", "rates-8"), "output": None},
    {"id": "run-72", "status": "committed", "inputs": ("orders-48", "rates-8"), "output": "revenue-49"},
]

def published_parents(releases, output_generation):
    matches = [run for run in releases if run["output"] == output_generation
               and run["status"] == "committed"]
    if len(matches) != 1:
        raise ValueError("output has no unique committed run")
    return matches[0]["inputs"]

assert published_parents(runs, "revenue-49") == ("orders-48", "rates-8")

Performance and operating cost

The direct lookup above scans R runs in O(R) time and stores O(R) records. Indexing committed output generation reduces expected lookup to O(1), while recording partition or column edges increases metadata volume. Retain evidence at least as long as the snapshots and audit period it must explain; sampling away successful release events destroys the proof chain.

Common Mistakes

  • Do not label a failed candidate as the parent of a published snapshot.
  • Do not infer all-to-all edges from a multi-input, multi-output job.
  • Do not identify a run solely by its scheduled timestamp.

Read next

Continue the workflow: Correction-aware quality alerts.

Continue the workflow: Lineage metadata minimization and access.

ai-data
data-engineering
Storage details