Plan a bounded delete-file maintenance run, prove current-row equivalence and keep a retained snapshot readable after publication.
Project: reduce delete overhead without breaking old snapshots
Create the workload
Load 470 order rows into three partitions. Mark 73 rows as deleted through several small delete batches, then create a retained snapshot before maintenance. Give one partition most of the delete files and most of the reader traffic. Record per-partition file counts, scan bytes, query latency and visible totals. The priority score should select the hot partition first rather than the largest partition by bytes.
Build a bounded plan
Set a daily rewrite ceiling of 768 MB and list the selected files, applicable delete generations and expected bytes written. Choose whether merging delete metadata or rewriting data files offers the better read-cost reduction. Name the snapshot and input sequence in the plan. A plan with no budget ceiling can consume capacity needed by ingestion or interactive reads.
Inject a concurrent update
Start rewriting from snapshot 74, then commit one new order and one delete to the same partition. The first candidate must fail conflict validation or be safely rebased. Retry with the new snapshot and prove that the newer insert is visible while the deleted old row remains absent. Snapshot safety is tested at both the current and retained generations.
Validate publication and cleanup
Compare current key sets, duplicate counts and amount totals with a pinned pre-rewrite oracle. Query the retained snapshot to show its older answer is still readable. Commit the replacement atomically. Run garbage-collection eligibility checks but leave old files in place while the replay horizon references them; record this as retained storage, not an incomplete maintenance commit.
Deliver the evidence
Provide the before-and-after file inventory, selection score, I/O budget, conflict trace, retry snapshot, two query outputs, p95 scan latency and cleanup eligibility report. Include a deliberately invalid candidate that omits one delete and show the gate rejecting it. A lower file count without row-equivalence proof is not a successful project result.
Implementation
source = {f"order-{order_number}": order_number * 25
for order_number in range(47, 517)}
deleted = {f"order-{order_number}" for order_number in range(47, 120)}
def current_snapshot(rows, removed):
return {key: amount for key, amount in rows.items() if key not in removed}
before = current_snapshot(source, deleted)
after = current_snapshot(source, deleted)
assert len(source) == 470
assert len(before) == 397
assert before == afterPerformance and operating cost
The reference comparison is O(N) expected time and O(N) state for N orders. The lakehouse exercise also incurs object opens, delete-file scans, candidate writes and one retained snapshot. A bounded maintenance plan controls compute spend; it does not authorize early deletion of files needed by readers pinned to older generations.
Common Mistakes
- Do not select work by file size alone when reader traffic differs.
- Do not publish a candidate that ignores a concurrent commit.
- Do not treat a correct current query as proof that retained snapshots survived.
