Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Delete-file amplification and compaction

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A lakehouse delete can avoid rewriting a large data file immediately, but accumulated delete metadata raises planning and read cost until maintenance catches up.

Distinguish logical removal from rewritten data

A reader may combine live data files with position deletes, equality deletes or a table-specific deletion bitmap. The row disappears from the current query, yet the original bytes can remain in older files and snapshots. Do not treat a successful row-level query as physical erasure. Erasure proof must examine retention and every copy separately.

Count work at the query boundary

One tiny delete file per transaction can make a scan open hundreds of objects before producing a few rows. Track applicable delete-file count, delete bytes, changed-row density, data-file size and query latency by partition. A table-wide average hides a hot partition. Run the same query before and after maintenance; file count alone is a proxy, not the user-facing outcome.

Choose the narrowest rewrite

If the data files are well sized but delete files are fragmented, compact compatible delete files without rewriting unaffected data. If many rows in a data file are deleted, rewrite that data file with surviving rows and publish a new snapshot. Preserve the delete applicability rules and sequence ordering; combining incompatible generations can resurrect a row or erase a newer insert.

Schedule by benefit and budget

Estimate bytes read, bytes written, metadata operations and expected scan reduction for each candidate partition. Prefer the hottest high-amplification partitions under a daily I/O budget. A low-traffic historical partition may safely wait. Coordinate maintenance with ingestion so repeated rewrites do not chase each new small batch; retention policy determines when old files may finally be removed.

Prove the post-commit state

Commit replacement files and metadata as one snapshot, then compare visible keys and control totals before moving consumers. A failed rewrite must leave the old snapshot queryable. Check readers pinned to older snapshots before deleting physical files. Measure query p95 after the commit and record whether the predicted benefit materialized; otherwise the scheduler may spend more than it saves.

Implementation

python
partitions = [
    {"name": "orders-west", "data_mb": 384, "delete_files": 47, "reads_per_day": 260},
    {"name": "orders-archive", "data_mb": 512, "delete_files": 6, "reads_per_day": 2},
]

def maintenance_priority(partition):
    return partition["delete_files"] * partition["reads_per_day"] / partition["data_mb"]

ranked = sorted(partitions, key=maintenance_priority, reverse=True)
assert ranked[0]["name"] == "orders-west"
assert maintenance_priority(ranked[0]) > maintenance_priority(ranked[1])

Performance and operating cost

Ranking P partitions costs O(P log P) time and O(P) space in this reference. Actual rewrite cost is dominated by bytes read and written, while every retained snapshot adds metadata and storage pressure. A cheap delete-file merge can reduce open overhead without reclaiming deleted row bytes; a data-file rewrite costs more but may reduce both scan work and later delete handling.

Common Mistakes

  • Do not infer physical erasure from a current-snapshot query.
  • Do not merge delete metadata across incompatible sequence boundaries.
  • Do not rewrite every partition on the same fixed schedule without measuring read benefit.

Read next

ai-data
data-engineering
Storage details