Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pretraining corpus boundaries and provenance audit

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Self-supervised training uses unlabeled records, but identities, periods and governance rules still determine whether the downstream evaluation is valid.

Freeze the test identity list first

A later parcel test set must be chosen before pretraining corpus selection. If the same physical parcel contributes unlabeled images to pretraining, an identity-sensitive encoder may appear to transfer while partly recognizing that parcel. Whether transductive exposure is allowed depends on the intended deployment claim; disclose it and use a stricter inductive test for new-parcel claims. Overlap auditing formalizes the check.

Record more than a file hash

A resized image and its original can have different byte hashes. Track parcel ID, shipment ID, capture timestamp, depot, camera and derived-image lineage. Exact hashes still catch duplicate files; entity keys catch harder overlaps. The code checks known IDs only and should be paired with near-duplicate image review.

Keep every processing fit inside its boundary

Pixel normalization, tokenizer construction, vocabulary statistics and augmentation calibration can all learn from a corpus. State which split was used to fit each transform. If final-test images tune transforms, freeze a new untouched test or qualify the claim. Validation design applies before model fitting, not after it.

Respect availability and retention

Unlabeled production images can contain personal data, expired retention records or licensed material. The model card should record source, usage rights, deletion handling and derived-artifact policy. These checks are operational requirements even when the objective uses no human labels.

Carry the manifest to release

Version the source query, exclusion rules, identity keys, transform settings and encoder checkpoint. A downstream score without this lineage cannot be reproduced or compared when the pool changes. Hold release if the overlap count or legal scope is unresolved. The release packet requires a clean manifest.

Implementation

python
pretraining_rows = [
    {"asset": "image-47a", "parcel_id": "dock-47", "captured_day": 3},
    {"asset": "image-62a", "parcel_id": "dock-62", "captured_day": 4},
    {"asset": "image-83a", "parcel_id": "dock-83", "captured_day": 6},
]
held_out_parcels = {"dock-83", "dock-94"}

def audit_known_overlap(rows, held_out_ids):
    return sorted({row["parcel_id"] for row in rows
                   if row["parcel_id"] in held_out_ids})

overlap = audit_known_overlap(pretraining_rows, held_out_parcels)
assert overlap == ["dock-83"]
clean_corpus = [row for row in pretraining_rows
                if row["parcel_id"] not in held_out_parcels]
assert len(clean_corpus) == 2

Performance and operating cost

With a hashed test-ID set, checking N pretraining records costs O(N) expected time and O(T) memory for T held-out identities. Near-duplicate matching costs more and may need embedding retrieval plus manual adjudication. Persisting lineage also adds storage, but prevents ambiguous evaluation claims.

Common Mistakes

  • Do not use different bytes as proof that two images show different parcels.
  • Do not fit preprocessing on final-test records unnoticed.
  • Do not call a corpus reusable without documenting its rights and retention scope.

Read next

ai-data
machine-learning
Storage details