Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Table snapshots and atomic publication

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A table snapshot names one committed set of data and delete files, letting readers observe a complete version while writers prepare the next.

Treat files as candidates until commit

An order correction may write three new files and retire two old ones. If a reader lists the directory halfway through, it can combine incompatible generations. Instead, write immutable candidate files, validate their row counts and checksums, then commit a new metadata pointer that names exactly the accepted files. Backfill publication follows the same boundary.

Understand conflict detection

Two writers can start from snapshot 41 and both prepare updates. A commit protocol must validate whether either writer changed data that the other relied on, then reject or retry a conflicting commit. Blindly replacing the pointer with the last writer's manifest loses earlier work. A retry must recalculate against the current snapshot rather than reuse an invalid old read set.

Keep reads pinned

A report that reads several files should pin a snapshot ID for its whole job. It must not fetch the newest pointer between partitions. Record the snapshot ID with the report and its query version so a later discrepancy can be reproduced. Analysis snapshots extend this discipline to code and cutoff time.

Account for row-level changes

Some table formats represent deletes separately from data files. The reader must apply all delete files valid for its snapshot; a raw Parquet reader that ignores them can expose deleted rows. Deletion file format and applicability vary by table specification, so test through the production engine before claiming the row is gone. Erasure verification must check every reader path.

Publish only after checks

For a daily order table, compare source count, accepted count, quarantined count, and distinct order keys before commit. If 47 source records become 45 accepted plus two quarantined, that balance is coherent; 44 plus two is not. A failed validation should leave the previous snapshot visible and the candidate files eligible for safe cleanup.

Implementation

python
snapshots = {41: {"orders-a", "orders-b"}}

def commit_snapshot(history, expected_parent, add_files, remove_files):
    current = max(history)
    if current != expected_parent:
        raise RuntimeError("stale parent snapshot")
    old_files = history[current]
    if not set(remove_files).issubset(old_files):
        raise ValueError("remove set not in parent")
    next_id = current + 1
    history[next_id] = (old_files - set(remove_files)) | set(add_files)
    return next_id

assert commit_snapshot(snapshots, 41, {"orders-c"}, {"orders-a"}) == 42
assert snapshots[41] == {"orders-a", "orders-b"}
assert snapshots[42] == {"orders-b", "orders-c"}

Performance and operating cost

The set operation shown costs O(F) time and space for F files in the parent manifest. Real table formats use metadata trees and optimistic commits to avoid copying every file name for every write. Snapshot metadata, conflict retries, and old-file retention add cost in exchange for consistent reads.

Common Mistakes

  • Do not expose candidate files through directory listing before commit.
  • Do not retry a conflicting write without rechecking its read set.
  • Do not read raw data files while ignoring snapshot-specific deletes.

Read next

Continue the workflow: Task retries and atomic partition output.

Continue the workflow: Optimistic table commits and write conflicts.

ai-data
data-engineering
Storage details