An optimistic table writer reads a snapshot, prepares new files and attempts an atomic metadata commit; a competing commit may force validation or retry.
Optimistic table commits and write conflicts
Name the read and write sets
A nightly account correction reads account-day rows and replaces the affected partition. A second writer appends new account-day rows. Whether they conflict depends on table isolation and which files or predicates each writer touched. Recording only the target table name is too coarse. Atomic publication makes each successful version visible at once, but does not settle conflicting intentions.
Do not blindly replay a failed commit
If writer A validates against snapshot 47 while writer B commits snapshot 48 first, A may need to recompute against 48. Reusing A’s old candidate files without checking the new table state can discard B’s rows or violate uniqueness. Distinguish a safe metadata retry from a business transformation that must be rerun. Pin source inputs and record the base snapshot.
Keep attempts off the reader path
Candidate files are not published merely because they exist. The commit points readers to one accepted metadata generation. An abandoned writer can leave unreferenced files; remove them only after in-flight attempts and rollback windows are ruled out. Task attempt selection handles this lower boundary.
Choose a conflict policy by grain
Two independent appends to disjoint partitions may both succeed after revalidation. Two updates to the same order version cannot be merged by appending both unless the product explicitly keeps both versions. Define unique keys and affected predicates, then fail or serialize competing changes that alter the same logical row. A retry that silently changes which correction wins is a semantic decision.
Exercise concurrent release
Start two writers from the same snapshot. Commit the first, force the second to validate against the new metadata, and assert either a correct rebase or an explicit conflict. Verify only one current snapshot pointer, complete files, unchanged unrelated partitions and a manifest naming both attempts. Measure orphan cleanup separately from commit success.
Implementation
table = {"version": 47, "rows": {"ord-47": 4200}}
def commit_update(current, base_version, order_id, cents):
if current["version"] != base_version:
raise RuntimeError("snapshot changed; recompute before commit")
next_rows = dict(current["rows"])
next_rows[order_id] = cents
return {"version": base_version + 1, "rows": next_rows}
first = commit_update(table, 47, "ord-48", 1900)
try:
commit_update(first, 47, "ord-47", 4300)
except RuntimeError:
pass
else:
raise AssertionError("stale writer committed")
assert first["rows"] == {"ord-47": 4200, "ord-48": 1900}Performance and operating cost
The reference copy costs O(K) time and space for K rows; real table formats avoid copying all data by writing changed files and a new metadata tree. Conflict validation adds metadata reads and may force repeated compute. Under contention, the cost of discarded candidate files and recomputation can dominate the small atomic pointer swap.
Common Mistakes
- Do not treat prepared files as committed rows.
- Do not retry a stale business transform without revalidating its read set.
- Do not delete abandoned files while another writer can still reference them.
Read next
- Table snapshots and atomic publication
- Task retries and atomic partition output
- Compaction, retention and the replay horizon
- Project: migrate an order event contract safely
- Replay manifests and audit trails
Continue the workflow: Snapshot-safe delete rewrites.
