A lakehouse delete can avoid rewriting a large data file immediately, but accumulated delete metadata raises planning and read cost until maintenance catches up.
Delete-file amplification and compaction
Distinguish logical removal from rewritten data
A reader may combine live data files with position deletes, equality deletes or a table-specific deletion bitmap. The row disappears from the current query, yet the original bytes can remain in older files and snapshots. Do not treat a successful row-level query as physical erasure. Erasure proof must examine retention and every copy separately.
Count work at the query boundary
One tiny delete file per transaction can make a scan open hundreds of objects before producing a few rows. Track applicable delete-file count, delete bytes, changed-row density, data-file size and query latency by partition. A table-wide average hides a hot partition. Run the same query before and after maintenance; file count alone is a proxy, not the user-facing outcome.
Choose the narrowest rewrite
If the data files are well sized but delete files are fragmented, compact compatible delete files without rewriting unaffected data. If many rows in a data file are deleted, rewrite that data file with surviving rows and publish a new snapshot. Preserve the delete applicability rules and sequence ordering; combining incompatible generations can resurrect a row or erase a newer insert.
Schedule by benefit and budget
Estimate bytes read, bytes written, metadata operations and expected scan reduction for each candidate partition. Prefer the hottest high-amplification partitions under a daily I/O budget. A low-traffic historical partition may safely wait. Coordinate maintenance with ingestion so repeated rewrites do not chase each new small batch; retention policy determines when old files may finally be removed.
Prove the post-commit state
Commit replacement files and metadata as one snapshot, then compare visible keys and control totals before moving consumers. A failed rewrite must leave the old snapshot queryable. Check readers pinned to older snapshots before deleting physical files. Measure query p95 after the commit and record whether the predicted benefit materialized; otherwise the scheduler may spend more than it saves.
Implementation
partitions = [
{"name": "orders-west", "data_mb": 384, "delete_files": 47, "reads_per_day": 260},
{"name": "orders-archive", "data_mb": 512, "delete_files": 6, "reads_per_day": 2},
]
def maintenance_priority(partition):
return partition["delete_files"] * partition["reads_per_day"] / partition["data_mb"]
ranked = sorted(partitions, key=maintenance_priority, reverse=True)
assert ranked[0]["name"] == "orders-west"
assert maintenance_priority(ranked[0]) > maintenance_priority(ranked[1])Performance and operating cost
Ranking P partitions costs O(P log P) time and O(P) space in this reference. Actual rewrite cost is dominated by bytes read and written, while every retained snapshot adds metadata and storage pressure. A cheap delete-file merge can reduce open overhead without reclaiming deleted row bytes; a data-file rewrite costs more but may reduce both scan work and later delete handling.
Common Mistakes
- Do not infer physical erasure from a current-snapshot query.
- Do not merge delete metadata across incompatible sequence boundaries.
- Do not rewrite every partition on the same fixed schedule without measuring read benefit.
