A grid key makes location data easier to prune, but choosing one cell resolution for every city can replace a full scan with thousands of tiny files and one overloaded downtown cell.
Spatial cell resolution and hotspot layout
Start with the question, not the grid
Map the actual predicates: a city dashboard reads days of events inside several neighborhoods, while a proximity alert reads a small radius around one point. Keep event time as an independent partition key because time filters usually eliminate more history than a spatial key. Add a coarse cell identifier for physical ordering or clustering, then measure files and bytes scanned for the real query shapes. Partition planning explains why a column is not automatically a good directory boundary.
Set resolution against file size
A finer grid yields tighter candidate sets, but it also increases partition count and small-file pressure. Count events, bytes and reader frequency per cell at several resolutions. The right resolution is one that prunes enough irrelevant data while leaving each active cell with files large enough for efficient scans. Keep sparse cells together where the table format allows it; do not create a directory for every empty possible cell.
Plan for uneven density
A downtown cell can receive 47 times the traffic of a rural cell. Splitting every cell more finely to fix one hotspot wastes metadata elsewhere. Use time slices, bounded salt buckets or an adaptive subdivision for the hot area, but document how a reader expands that layout back into a complete candidate set. Skew handling also matters when events are shuffled for downstream aggregation.
Keep the exact geometry
A cell identifier is an index candidate, not the location itself. A radius can cross cell edges; a polygon can touch many cells without containing every point in them. Retain the original coordinates and coordinate reference assumptions, select all potentially intersecting cells and apply an exact point-in-polygon or distance predicate afterward. The coarse filter may over-read, but it must never discard a true match.
Check the release boundary
When a location correction moves an event from one cell to another, the old placement must disappear from the current generation before the new placement is published. Compare stable event IDs, cell counts and query results across the two generations. A single corrected point can otherwise appear twice near a boundary. Keep the older snapshot for the replay period and record which grid version produced each published key.
Implementation
from collections import Counter
traffic_by_cell = Counter({"downtown-47": 47000, "suburb-48": 960,
"rural-49": 125})
target_events_per_file = 8000
def planned_files(events_per_cell, target_size):
return {cell: max(1, (events + target_size - 1) // target_size)
for cell, events in events_per_cell.items()}
plan = planned_files(traffic_by_cell, target_events_per_file)
assert plan == {"downtown-47": 6, "suburb-48": 1, "rural-49": 1}
assert sum(traffic_by_cell.values()) == 48085Performance and operating cost
Counting N events by cell takes O(N) expected time and O(C) state for C populated cells. Planning files from those counts is O(C). Real query cost also includes metadata listing and scanning every selected file; excessively fine cells raise both file count and compaction work. Compare end-to-end latency and bytes scanned, not cell count alone.
Common Mistakes
- Do not treat a grid cell as an exact distance or polygon answer.
- Do not partition so finely that most cells contain one tiny file.
- Do not repair a hotspot by changing the grid without a versioned migration plan.
