A deleted source record may survive in feature tables, snapshots, evaluation files and model artifacts unless lineage drives a policy-specific response.
MLOps retention: trace source deletion into features and training artifacts
Map data derivatives before deciding action
A receipt can appear in raw storage, cleaned tables, offline features, online feature caches, training snapshots, sampled debug logs and reports. A trained model may contain information from many records. Record lineage from source class to derivative artifact and the retention policy for each. A deletion request may require removing some derivatives, rebuilding others or preserving narrowly scoped evidence under a valid obligation. Do not promise that deleting one row automatically removes its influence from an existing model. Training manifests help locate affected snapshots.
Distinguish deletion from withdrawal
A pipeline may delete a source row while a published scoring output remains active. Withdraw or supersede operational outputs under their consumer contract, then clear caches and queued work where required. Keep a restricted tombstone or deletion receipt if policy allows, so a later backfill does not resurrect the record from an old snapshot. Batch replay can produce replacement outputs after source corrections, but customer actions need their own idempotent withdrawal rule.
Decide model-level treatment explicitly
If a training snapshot contained data that must no longer influence a model, assess whether the model must be retired, retrained from an eligible snapshot or handled under a documented exception. The answer depends on legal and product policy; a generic hash or feature deletion cannot remove learned parameters. Record affected model digests and owners. Retirement gates prevent deleting an affected model before traffic has moved to an approved replacement.
Verify the whole path
Exercise a test deletion with a source row present in an online cache, a batch output and a training snapshot. Check every derivative under the applicable policy, record unresolved items and test a backfill to ensure the row stays excluded. The retirement project keeps a minimal evidence ledger while removing unnecessary copies and replacing an affected model where required.
Implementation
def affected_derivatives(children_by_parent, source_id):
pending = [source_id]
seen = {source_id}
while pending:
parent = pending.pop()
for child in children_by_parent.get(parent, ()):
if child not in seen:
seen.add(child)
pending.append(child)
seen.remove(source_id)
return seen
lineage = {"receipt-47": {"feature-82", "batch-129"},
"feature-82": {"snapshot-91"},
"snapshot-91": {"model-r7"}}
assert affected_derivatives(lineage, "receipt-47") == {
"feature-82", "batch-129", "snapshot-91", "model-r7"}
Performance and operating cost
Traversing reachable lineage takes O(v + e) time and O(v) space for v visited artifacts and e edges. Large shared artifacts can make the affected set broad, so record artifact-level lineage and policy-specific actions. Reachability identifies candidates for review; it does not determine whether deleting, retraining or retaining each item is lawful or technically sufficient.
Common Mistakes
- Assuming raw-row deletion erases a model’s learned influence.
- Leaving an online feature cache or batch output active after source withdrawal.
- Allowing a historical backfill to resurrect a deleted record.
- Deleting audit evidence without checking its separate retention obligation.
Read next
- Model retirement: find consumers before removing a version
- Project: retire a receipt model without breaking batch or audit
- Training manifests: link data, code, configuration and artifact
- Batch replay: supersede outputs without duplicating downstream actions
- Inference logs: keep diagnostic joins without copying sensitive payloads
