Build a small external telemetry release that rejects incomplete files and stale metadata before an analyst query can use it.
Project: publish telemetry with a catalog freshness gate
Stage a measurable workload
Create 47 hourly telemetry partitions with immutable object names. Give one hour 23 files and the others three each, so a broad listing and a pruned query have visibly different planning costs. Record the source generation, object inventory and query predicate. Establish baseline planning time, listed-object count and scan bytes. Partition design affects the result independently of cache policy.
Publish one hour safely
Write a manifest for the next hour with expected paths, sizes, record counts and a batch generation. Omit one object on the first attempt. The gate must reject it before catalog registration and keep the last published hour readable. Add the missing object, rerun the gate, register the partition and then refresh the metadata cache. Store the completed refresh generation, not merely its start time.
Exercise a stale reader
Issue a query that requires the newly published generation while the cache still names the previous one. It must refresh, use a permitted direct listing or return a freshness error under the chosen contract. It must not return zero rows as though there were no telemetry. Cache freshness is checked at the query boundary, not inferred from the ingest success log.
Compare engines and corrections
Read the same published hour from a second engine. If one engine derives partitions from a pattern and the other uses registered entries, prove that both see the same manifest files. Publish a corrected generation with one changed record, and ensure readers move as a unit instead of reading old and new objects together. Leave retained files until their replay horizon closes.
Submit release evidence
Provide the failed and successful gate reports, inventory diff, catalog entry, cache generation, reader response, p95 planning time and request count. Include a trace where a table location changes but the cache has not refreshed; the gate must catch it. A fast query with stale or partial data fails this project even if its latency chart looks excellent.
Implementation
expected_files = {
"hour=09/telemetry-47.parquet": 470,
"hour=09/telemetry-48.parquet": 480,
}
observed_files = {"hour=09/telemetry-47.parquet": 470}
def can_publish(manifest_files, object_files, catalog_generation, requested_generation):
return (manifest_files == object_files
and catalog_generation == requested_generation)
assert not can_publish(expected_files, observed_files, "telemetry-46", "telemetry-47")
observed_files["hour=09/telemetry-48.parquet"] = 480
assert can_publish(expected_files, observed_files, "telemetry-47", "telemetry-47")Performance and operating cost
The reference gate compares F file entries in O(F) expected time and O(F) stored inventory. Production cost also includes object listings, catalog registration, cache refresh and two engine queries. A targeted prefix refresh saves work when arrivals are predictable. The safety gate may add latency to publication; measure it against the cost of serving a partial interval.
Common Mistakes
- Do not return a stale empty result for a newly published hour.
- Do not treat a catalog timestamp as proof that a refresh completed.
- Do not let two engines compare different file generations.
Read next
- External-table metadata cache and refresh
- Partition registration and manifest gates
- Replay manifests and audit trails
- Serving indexes and freshness contracts
- Project: migrate a dashboard with shadow reads
Continue the workflow: Project: publish a boundary-safe delivery-zone lake.
