Skip to content
AITroveRead. Build. Understand.
Make this comfortable

MLOps retention: trace source deletion into features and training artifacts

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A deleted source record may survive in feature tables, snapshots, evaluation files and model artifacts unless lineage drives a policy-specific response.

Map data derivatives before deciding action

A receipt can appear in raw storage, cleaned tables, offline features, online feature caches, training snapshots, sampled debug logs and reports. A trained model may contain information from many records. Record lineage from source class to derivative artifact and the retention policy for each. A deletion request may require removing some derivatives, rebuilding others or preserving narrowly scoped evidence under a valid obligation. Do not promise that deleting one row automatically removes its influence from an existing model. Training manifests help locate affected snapshots.

Distinguish deletion from withdrawal

A pipeline may delete a source row while a published scoring output remains active. Withdraw or supersede operational outputs under their consumer contract, then clear caches and queued work where required. Keep a restricted tombstone or deletion receipt if policy allows, so a later backfill does not resurrect the record from an old snapshot. Batch replay can produce replacement outputs after source corrections, but customer actions need their own idempotent withdrawal rule.

Decide model-level treatment explicitly

If a training snapshot contained data that must no longer influence a model, assess whether the model must be retired, retrained from an eligible snapshot or handled under a documented exception. The answer depends on legal and product policy; a generic hash or feature deletion cannot remove learned parameters. Record affected model digests and owners. Retirement gates prevent deleting an affected model before traffic has moved to an approved replacement.

Verify the whole path

Exercise a test deletion with a source row present in an online cache, a batch output and a training snapshot. Check every derivative under the applicable policy, record unresolved items and test a backfill to ensure the row stays excluded. The retirement project keeps a minimal evidence ledger while removing unnecessary copies and replacing an affected model where required.

Implementation

python
def affected_derivatives(children_by_parent, source_id):
    pending = [source_id]
    seen = {source_id}
    while pending:
        parent = pending.pop()
        for child in children_by_parent.get(parent, ()):
            if child not in seen:
                seen.add(child)
                pending.append(child)
    seen.remove(source_id)
    return seen

lineage = {"receipt-47": {"feature-82", "batch-129"},
           "feature-82": {"snapshot-91"},
           "snapshot-91": {"model-r7"}}
assert affected_derivatives(lineage, "receipt-47") == {
    "feature-82", "batch-129", "snapshot-91", "model-r7"}

Performance and operating cost

Traversing reachable lineage takes O(v + e) time and O(v) space for v visited artifacts and e edges. Large shared artifacts can make the affected set broad, so record artifact-level lineage and policy-specific actions. Reachability identifies candidates for review; it does not determine whether deleting, retraining or retaining each item is lawful or technically sufficient.

Common Mistakes

  • Assuming raw-row deletion erases a model’s learned influence.
  • Leaving an online feature cache or batch output active after source withdrawal.
  • Allowing a historical backfill to resurrect a deleted record.
  • Deleting audit evidence without checking its separate retention obligation.

Read next

ai-data
mlops
Storage details