Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: prove erasure after a regional restore

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Release a restored customer mart only after regional placement, tombstone replay, version inventory and consumer-read checks agree.

Set the contract and inventory

A customer mart has a primary, one disaster-recovery replica, a serving index and a retained backup. Customer 47 is erased at source position 73, while the newest backup stops at 71. Classify each artifact, record its region and list which copies may legally exist. The regional gate should deny an unauthorized destination before any restore files are written.

Simulate an incomplete delete

Remove the current primary row but leave one older object version in the replica. A current-read test may pass while a version-specific read still retrieves data. The erasure workflow must report pending, not complete. Delete or expire the historical version when policy permits, then verify the version inventory again. If a hold blocks removal, record its authority and keep the result explicitly pending.

Restore behind a publication gate

Restore the backup to an authorized region into candidate storage, replay the deletion ledger from position 72 onward and build the serving index from the corrected snapshot. Inject an update at 72 after the delete at 73; source ordering must reject it. The restore guard must pass before the table pointer or index pointer changes.

Check all reader surfaces

Run current and version-specific reads against both object stores, query the restored mart and serving index, inspect dead-letter copies and confirm no old snapshot pointer is publicly selectable. Compare deletion-ledger position, table generation and index generation in one manifest. A missing response from one region is not a pass; it is an unverified result that blocks completion.

Deliver a defensible proof bundle

Provide the classification decision, copy inventory, denied egress attempt, deletion-ledger entry, candidate restore manifest, version listings, hold status, post-replay query output and publication timestamp. Include a rerun that shows the process is idempotent. Separate logical hiding, inaccessible retained bytes and verified physical removal in the final state report.

Implementation

python
copies = [
    {"region": "region-west", "version": "v-71", "customer": "customer-47", "present": False},
    {"region": "region-east", "version": "v-68", "customer": "customer-47", "present": True},
]

def erasure_complete(inventory, customer_id):
    return all(not item["present"] for item in inventory
               if item["customer"] == customer_id)

assert not erasure_complete(copies, "customer-47")
copies[1]["present"] = False
assert erasure_complete(copies, "customer-47")

Performance and operating cost

The inventory check is O(V) time over V recorded versions and O(1) auxiliary space. A real proof costs version listings, regional reads, snapshot scans and restore validation; it must include every known copy and detect inventory gaps. Candidate storage and index rebuild add temporary cost, but publishing first would expose deleted data again.

Common Mistakes

  • Do not declare completion from one region or one current-read response.
  • Do not publish a restored table before applying the later deletion ledger.
  • Do not hide a pending legal hold inside a successful erasure status.

Read next

ai-data
data-engineering
Storage details