Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Partition registration and manifest gates

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A partition becomes queryable only when its data, catalog entry and release marker agree; a directory name by itself proves little.

Choose a discovery model

Explicit registration lists only known partitions, while projection derives possible locations from a pattern. Projection avoids a per-partition catalog write, but it can invent empty locations and may be specific to one query engine. Explicit registration costs metadata operations and can lag ingestion. Choose based on partition count, sparsity, engine mix and freshness needs, then test another engine rather than assuming it reads the same projected inventory.

Make the release marker authoritative

Write all objects for a batch under immutable names and validate their schema, row count and checksums. Build a manifest naming the expected objects and their sizes. Only then register the partition or advance a published marker. A consumer should require a complete manifest before treating a newly visible directory as a finished interval. Atomic partition output still matters when a task retries and leaves abandoned objects.

Compare three inventories

At promotion, compare the manifest file set, object-store listing and catalog-visible partitions. The sets need not be identical at every instant, but the published files must all be present, the catalog must target the correct prefix and no unapproved file may be read. Check object size or checksum where available; matching file names alone will miss an overwritten object. Save the comparison report with the release generation.

Handle correction and removal

A corrected hour should publish a new immutable generation and switch the catalog or table pointer to it. Retire older objects only after reader and replay retention permits. Removing a partition entry before its replacement is ready creates a query gap; adding the replacement without excluding old files can double count. Retention rules determine when stale bytes may be collected.

Test the broken boundary

Inject one missing object, one stray object, a catalog entry pointing at the previous generation and a projection range that omits a valid date. The gate should reject each release with a concrete reason. Also show that a projected empty location is not an ingestion success signal. A second engine must resolve the same published file set before cross-engine comparison is meaningful.

Implementation

python
manifest = {"generation": "sensor-47", "files": {
    "day=27/hour=08/part-47.parquet": 2375,
    "day=27/hour=08/part-48.parquet": 6400,
}}
objects = dict(manifest["files"])
catalog_prefix = "day=27/hour=08/"

def release_issues(expected, observed, prefix):
    missing = set(expected) - set(observed)
    unexpected = set(observed) - set(expected)
    changed = {path for path in expected.keys() & observed.keys()
               if expected[path] != observed[path]}
    wrong_prefix = {path for path in expected if not path.startswith(prefix)}
    return missing, unexpected, changed, wrong_prefix

assert release_issues(manifest["files"], objects, catalog_prefix) == (set(), set(), set(), set())
assert release_issues(manifest["files"], {"day=27/hour=08/part-47.parquet": 2375}, catalog_prefix)[0]

Performance and operating cost

Set comparison is O(F) expected time and O(F) auxiliary space for F files. Full object inventory and checksums can dominate wall time and request spend; compare only changed prefixes when the storage and publication contract allows it. Projection may save catalog writes but can increase planning work for many nonexistent partitions.

Common Mistakes

  • Do not register a partition before every expected object is durable.
  • Do not mistake a projected location for a completed batch.
  • Do not replace objects in place while readers or caches retain old metadata.

Read next

ai-data
data-engineering
Storage details