Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Deletion propagation and erasure proof

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A deletion request is complete only when every authorized serving path stops returning the subject and the remaining retained copies follow a defined erasure schedule.

Inventory the copies

One customer identifier may appear in a raw event log, current customer table, warehouse history, feature snapshot, search index, export, and backup. Deleting only the primary table leaves other paths readable. Maintain a dataset inventory with owner, purpose, key format, retention, deletion method, and verification query. Lineage helps find descendants, but untracked manual exports need a separate inventory.

Distinguish logical and physical removal

A table delete may mark rows invisible in the current snapshot while their bytes remain in old files or snapshots until rewrite and retention cleanup. That can satisfy an immediate serving requirement, but physical erasure requires additional work. Define which guarantees apply to live queries, time travel, raw storage, caches, and backups; do not treat one SQL DELETE as proof for all of them.

Track an erasure job

Assign a request ID and normalized subject key, then record every target dataset, attempted operation, completion state, and evidence timestamp. Deduplicate repeated requests. A transient failure in one downstream index should remain visible and retryable, while already completed targets should not be reinserted from an old backfill. Backfill publication must consult the active deletion ledger.

Verify through production readers

Run the same queries and API reads that users and internal jobs use, pinned to the post-delete snapshot where applicable. Check the current table, historical views, and downstream serving stores. A count of zero in one warehouse table is insufficient if a cached endpoint still returns the customer. For backups that cannot be modified in place, record expiry and require reapplication of the ledger during restore.

Prevent reappearance

Keep tombstones or a suppression ledger long enough that delayed CDC, retries, and raw replay cannot recreate the subject. Treat a create event older than the deletion boundary as stale. If the business legitimately creates a new account with the same identifier, define a new entity epoch; otherwise key reuse can make the suppression rule ambiguous.

Implementation

python
deletion_ledger = {"cust-47": 131}
events = [
    {"customer": "cust-47", "source_position": 128, "kind": "upsert"},
    {"customer": "cust-83", "source_position": 129, "kind": "upsert"},
]

def publishable_events(records, tombstones):
    return [record for record in records
            if record["customer"] not in tombstones]

assert [record["customer"] for record in publishable_events(events, deletion_ledger)] == ["cust-83"]

Performance and operating cost

Filtering N events against a hashed ledger is O(N) expected time and O(N) output space; ledger storage grows with active suppression keys. Full physical erasure can require rewriting files and waiting through backup retention. The code demonstrates a replay barrier, not a substitute for authorization or legal retention policy.

Common Mistakes

  • Do not call a logical table delete physical erasure.
  • Do not let backfills ignore the deletion ledger.
  • Do not claim completion from one table query when caches or exports remain.

Read next

ai-data
data-engineering
Storage details