Rebuild a bounded order-lake partition with verified file bounds, benchmark it against the same snapshot and publish only a net-positive layout.
Project: release a measured order-lake layout
Construct the baseline
Generate 47,000 orders across 29 days and 83 accounts. Keep order IDs unique, add null customer regions and several wide text attributes, then write files with overlapping account ranges. Record committed snapshot ID, file count, min/max/null statistics and query workload. Include an account equality lookup, a day-and-account filter, a regional rollup and a full scan. The skip rules decide which files remain candidates.
Build a bounded candidate
Retain day as the partition key; sort recent high-traffic days by account ID within a bounded file-size target. Do not repartition to one directory per account. Regenerate statistics from candidate bytes and verify that all old rows remain represented under the table format’s delete semantics. A writer racing with the rewrite should trigger conflict handling or a new candidate rather than silently replacing its fresh rows.
Benchmark without moving the goalposts
Pin source and candidate snapshots that contain the same logical rows. Run each query with both cold and warm cache conditions, collect listed files, skipped files, row groups, bytes read and elapsed time, and repeat enough times to see variance. Measure candidate write time and storage amplification. The workload decision uses weighted costs, not one cherry-picked fast lookup.
Publish through a release gate
Compare row counts, checksums by day, delete behavior and query answers. Reject the candidate if any result differs, even if its skip ratio looks excellent. Commit the new file set atomically, retain the previous snapshot through the rollback window, and record file IDs and rewrite attempt in a manifest. Snapshot publication keeps readers on one complete generation.
Monitor what happens next
Append an intentionally unsorted day and show the skip ratio falling for that day. Derive a threshold for another bounded maintenance pass using read waste and rewrite cost. Submit the baseline and candidate manifests, benchmark table, validation output, conflict trace, release pointer and rollback test. A static layout screenshot is not evidence that the new writes remain efficient.
Implementation
file_ranges = [
("orders-a.parquet", 21, 39),
("orders-b.parquet", 40, 58),
("orders-c.parquet", 59, 83),
("orders-unknown.parquet", None, None),
]
def candidate_files(account_id, ranges):
return [file_name for file_name, minimum, maximum in ranges
if minimum is None or maximum is None
or minimum <= account_id <= maximum]
assert candidate_files(47, file_ranges) == [
"orders-b.parquet", "orders-unknown.parquet"]
baseline_order_ids = {"order-47", "order-61", "order-83"}
candidate_order_ids = {"order-83", "order-47", "order-61"}
assert candidate_order_ids == baseline_order_idsPerformance and operating cost
The file-metadata filter is O(F) for F files; a real rewrite sorts N selected records in roughly O(N log N) comparison time and needs shuffle or external-spill capacity. Candidate files temporarily duplicate storage until old snapshots expire. Report net read savings over the expected query volume against that rewrite, metadata, write-lag and storage cost.
Common Mistakes
- Do not release a candidate that changes query answers.
- Do not compute new bounds from a sample and then prune with them as if complete.
- Do not report performance without pricing the rewrite and unsorted future appends.
