File and row-group bounds can eliminate irrelevant reads, but missing or stale statistics must never eliminate a matching record.
Lake file statistics and safe data skipping
Start with the query predicate
A query for orders from account 47 can skip a file only when trustworthy statistics prove its account range excludes 47. A file with minimum 40 and maximum 60 remains a candidate even if it contains no 47; bounds allow false positives, not false negatives. Null predicates need their own null counts or explicit uncertainty. Row-group and column pruning happens at a different layer from skipping entire files.
Bind bounds to the committed file
Collect minimum, maximum, null count and record count from the actual written data, then attach them to the file version referenced by the committed snapshot. A rewritten file needs new statistics; copying old bounds to new content is unsafe. Files whose statistics are absent, truncated or computed with incompatible comparison rules remain candidates. A scan engine may still choose to read more than the minimum safe set for planning reasons.
Respect types and comparison rules
String ordering, decimal scale, timezone normalization and case folding can alter how a predicate compares with stored bounds. Treat a type migration as a statistics-contract change. If a query casts an account ID to text while the file stores an integer range, do not apply integer pruning unless the planner proves equivalence. Cross-engine semantic differences make a plan from one engine insufficient proof for another.
Measure actual pruning
Record files listed, files rejected by metadata, row groups read, bytes fetched and query time for a fixed workload. A file-level skip ratio of 90 percent may still be poor if the remaining files are wide or contain most rows. Compare cold and warm runs separately, and pin the same snapshot for baseline and candidate so concurrent appends cannot manufacture a gain. Missing statistics should be visible as a reason, not folded into a generic scanned count.
Keep deletion semantics intact
A file that carries a matching value may still have those rows removed by delete files. Pruning must use the table format’s current snapshot and delete semantics, not a loose directory scan of physical objects. Rewrites and delete materialization can alter bounds. Snapshot-safe rewrite governs publication; the skip checker should validate candidate output against a full scan before promotion.
Implementation
def file_may_match(file_stats, lower_account, upper_account):
minimum = file_stats.get("min_account")
maximum = file_stats.get("max_account")
if minimum is None or maximum is None:
return True # Unknown bounds must remain in the scan.
return not (maximum < lower_account or minimum > upper_account)
committed_files = [
{"file": "orders-47.parquet", "min_account": 41, "max_account": 53},
{"file": "orders-83.parquet", "min_account": 71, "max_account": 91},
{"file": "orders-61.parquet"},
]
candidates = [item["file"] for item in committed_files
if file_may_match(item, 47, 47)]
assert candidates == ["orders-47.parquet", "orders-61.parquet"]Performance and operating cost
Testing F file metadata records costs O(F) time and O(C) output space for C candidates; manifest indexes can reduce planning work for selective partitions. Statistical collection adds write CPU and metadata bytes. A false positive costs an extra read, while a false negative loses data, so an uncertain bound must default to scanning.
Common Mistakes
- Do not discard a file merely because it lacks statistics.
- Do not reuse pre-rewrite bounds for newly written bytes.
- Do not interpret a high skip ratio as a correctness test.
