Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Compaction, retention and the replay horizon

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Compaction rewrites small or fragmented files into fewer files; retention decides when old snapshots and their unreferenced files may be removed.

Name the reason to compact

A stream writing 2,400 five-kilobyte files per hour can spend more time opening files than scanning rows. Bin-packing them into larger files lowers open and planning overhead. It does not reduce the logical row count or repair duplicate records. Measure file count, average compressed size, query planning time, and bytes rewritten before scheduling the job.

Separate a rewrite from publication

A compactor writes replacement files, validates row totals and key-level checksums, and commits a new snapshot that swaps file references. Readers pinned to the older snapshot keep seeing the old files. A failed rewrite must leave old references intact. Snapshot commits provide this isolation; deleting files as soon as the replacement is written does not.

Set retention from real dependencies

A month-end audit may need a snapshot from 35 days ago even when routine dashboards need only a week. A delayed CDC replay may need source logs, original raw records, schemas, and the table snapshot that formed its input. List each recovery and audit requirement, then set retention to the longest required horizon plus operational margin. CDC replay fails if its earliest necessary source position has expired.

Prove cleanup safety

Old snapshots can reference files still needed by pinned readers, rollback plans, or branches. First expire eligible references according to the table's retention rules; then remove only orphan files that no live reference can reach, after a delay that covers in-flight writes. Keep a dry-run manifest and deletion count. Mistaking a candidate write for an orphan is a data-loss event.

Watch write amplification

Compacting every partition on every ingest batch repeatedly rewrites the same bytes. Target partitions whose file count or average size crosses a measured threshold. A hot partition may need scheduled compaction after its late-arrival window closes. Skew diagnostics identify partitions that deserve separate treatment.

Implementation

python
files = [
    {"name": "day47-a", "day": 47, "bytes": 28_000},
    {"name": "day47-b", "day": 47, "bytes": 34_000},
    {"name": "day48-a", "day": 48, "bytes": 156_000},
]

def compaction_candidates(manifest, min_file_bytes=50_000):
    by_day = {}
    for item in manifest:
        if item["bytes"] < min_file_bytes:
            by_day.setdefault(item["day"], []).append(item["name"])
    return {day: names for day, names in by_day.items() if len(names) > 1}

assert compaction_candidates(files) == {47: ["day47-a", "day47-b"]}

Performance and operating cost

Scanning F file summaries is O(F) time and O(P + C) memory for P candidate partitions and C candidate files. A real rewrite costs read and write I/O for every selected byte. Excessive compaction trades query savings for higher storage and compute bills.

Common Mistakes

  • Do not delete old files before all retained snapshots stop referring to them.
  • Do not compact on a timer without checking file sizes and write amplification.
  • Do not set retention shorter than replay, audit, or rollback requirements.

Read next

ai-data
data-engineering
Storage details