Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pretraining overlap and provenance audit

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A transferred model needs a record of checkpoint origin and known training overlap so target evaluation does not quietly reuse the same units or restricted data.

Trace the checkpoint

Store the source model identifier, license terms, intended use, source-data description, preprocessing version and model artifact digest. If the source training set is opaque, say so; there is no general way to prove non-overlap from a checkpoint alone. The code compares known parcel group IDs in a source manifest with target evaluation groups. The encoder contract pins the artifact being audited.

Check groups, not only exact files

A source photo and target test photo of the same parcel can have different bytes after cropping, compression or a new camera angle. Exact hashes catch exact duplicates but not that shared evaluation unit. Check parcel, capture session and incident group IDs where available; use visual similarity review for uncertain cases. Augmentation leakage explains why variants belong together.

Separate declared absence from unknown provenance

A zero intersection between two known manifests means those manifests share no listed IDs. It does not establish that the source model never saw the target photos if its full pretraining corpus is undisclosed. Record the scope of the check and hold out a newly collected target period where possible. Do not manufacture a clean-provenance claim from missing metadata. Target grouping still matters when the source corpus is unknown.

Review rights and data handling

Checkpoint reuse may be limited by license, collection consent, retention policy or security review. Keep these checks in the editorial and engineering workflow; this tutorial does not grant rights to any third-party model or dataset. If the allowed use is uncertain, resolve it before deployment, even if target validation looks strong. Image consent supplies the data-side question.

Version the audit with each update

A later checkpoint or a refreshed pretraining set changes the overlap boundary. Attach the audit to the exact artifact digest and target evaluation cohort, then rerun it when either changes. The release review treats unresolved provenance as a decision blocker.

Implementation

python
import hashlib

source_groups = {"parcel-A17", "parcel-B24", "parcel-C05"}
target_test_groups = {"parcel-C05", "parcel-D91", "parcel-E44"}
checkpoint_bytes = b"parcel-encoder-release-4"

def provenance_audit(known_source_groups, evaluation_groups, artifact_bytes):
    overlap = sorted(known_source_groups & evaluation_groups)
    return {
        "known_group_overlap": overlap,
        "source_groups_known": len(known_source_groups),
        "target_groups_checked": len(evaluation_groups),
        "artifact_sha256": hashlib.sha256(artifact_bytes).hexdigest(),
    }

audit = provenance_audit(source_groups, target_test_groups, checkpoint_bytes)
assert audit["known_group_overlap"] == ["parcel-C05"]
assert len(audit["artifact_sha256"]) == 64

Performance and operating cost

Intersecting S source and T target group IDs costs O(S + T) expected time and O(S + T) set memory; hashing an artifact costs O(B) time for B bytes. Manual near-duplicate review and verifying legal or consent conditions may be more expensive than either operation.

Common Mistakes

  • Do not claim no source overlap when the source training manifest is unavailable.
  • Do not rely on byte equality to detect all parcel-level duplicates.
  • Do not transfer a checkpoint without recording its permitted use and exact version.

Read next

Continue the workflow: Pretraining corpus boundaries and provenance audit.

ai-data
machine-learning
Storage details