Parquet groups rows into row groups and stores each column's values together, so a query can skip unneeded columns and sometimes whole row groups.
Parquet row groups, projection and scan cost
Connect layout to a query
An order analyst selecting order ID, status, and amount from a 63-column table should not scan every column chunk. Column projection avoids irrelevant bytes. If a row group has useful min and max statistics for event date, a date filter may skip that group; the engine must still read a group whose range overlaps the filter. Statistics are an optimization, not a correctness filter.
Distinguish file from table
A Parquet file contains data and footer metadata. It does not, by itself, supply a multi-file transaction, snapshot isolation, or a reliable table-wide deletion protocol. Use a table format or a controlled manifest when readers need a consistent set of files. Snapshot publication makes that distinction concrete.
Plan row groups and files together
A very small row group has high metadata overhead and weak scan amortization. A huge row group may waste reads when predicates are selective and can limit parallel work. Many tiny files increase listing, open, planning, and task-launch costs even if the compressed bytes are few. Choose sizes from representative queries and measured storage behavior rather than one universal target.
Sort for the filters that matter
Grouping nearby event dates can make row-group ranges tighter and improve skipping for time-bounded queries. Sorting on every candidate column is impossible; it adds write work and may help one query while hurting another. Compare bytes scanned, files opened, and elapsed time before and after a layout change. Partition design determines a coarser pruning boundary.
Watch evolving schemas
Files written under several schemas can coexist. Readers must resolve missing columns, field types, and defaults consistently across them. A read that silently converts missing amount to zero can create false revenue. The schema contract should state whether the field is absent, null, or truly zero.
Implementation
row_group_stats = [
{"file": "orders-a", "min_day": 40, "max_day": 46, "bytes": 120_000},
{"file": "orders-b", "min_day": 47, "max_day": 53, "bytes": 145_000},
{"file": "orders-c", "min_day": 54, "max_day": 60, "bytes": 132_000},
]
def candidate_groups(groups, first_day, last_day):
return [group for group in groups
if group["max_day"] >= first_day and group["min_day"] <= last_day]
scan = candidate_groups(row_group_stats, 48, 50)
assert [group["file"] for group in scan] == ["orders-b"]
assert sum(group["bytes"] for group in scan) == 145_000Performance and operating cost
Planning from G row-group statistics takes O(G) work in this example; a table catalog may prune earlier. Projection reduces bytes for unused columns, while row-group skipping reduces candidate groups. Neither eliminates a costly join or an incorrect table grain.
Common Mistakes
- Do not call a Parquet file an ACID table.
- Do not expect min/max skipping when ranges overlap widely or statistics are absent.
- Do not optimize compressed bytes alone while file-open overhead dominates.
Read next
- Partition planning, pruning and skew
- Table snapshots and atomic publication
- Compaction, retention and the replay horizon
- Schema compatibility and consumer rollout
- Project: release a versioned order lake
Continue the workflow: Task retries and atomic partition output.
Continue the workflow: Pipeline cost attribution and right-sizing.
