A transferred model needs a record of checkpoint origin and known training overlap so target evaluation does not quietly reuse the same units or restricted data.
Pretraining overlap and provenance audit
Trace the checkpoint
Store the source model identifier, license terms, intended use, source-data description, preprocessing version and model artifact digest. If the source training set is opaque, say so; there is no general way to prove non-overlap from a checkpoint alone. The code compares known parcel group IDs in a source manifest with target evaluation groups. The encoder contract pins the artifact being audited.
Check groups, not only exact files
A source photo and target test photo of the same parcel can have different bytes after cropping, compression or a new camera angle. Exact hashes catch exact duplicates but not that shared evaluation unit. Check parcel, capture session and incident group IDs where available; use visual similarity review for uncertain cases. Augmentation leakage explains why variants belong together.
Separate declared absence from unknown provenance
A zero intersection between two known manifests means those manifests share no listed IDs. It does not establish that the source model never saw the target photos if its full pretraining corpus is undisclosed. Record the scope of the check and hold out a newly collected target period where possible. Do not manufacture a clean-provenance claim from missing metadata. Target grouping still matters when the source corpus is unknown.
Review rights and data handling
Checkpoint reuse may be limited by license, collection consent, retention policy or security review. Keep these checks in the editorial and engineering workflow; this tutorial does not grant rights to any third-party model or dataset. If the allowed use is uncertain, resolve it before deployment, even if target validation looks strong. Image consent supplies the data-side question.
Version the audit with each update
A later checkpoint or a refreshed pretraining set changes the overlap boundary. Attach the audit to the exact artifact digest and target evaluation cohort, then rerun it when either changes. The release review treats unresolved provenance as a decision blocker.
Implementation
import hashlib
source_groups = {"parcel-A17", "parcel-B24", "parcel-C05"}
target_test_groups = {"parcel-C05", "parcel-D91", "parcel-E44"}
checkpoint_bytes = b"parcel-encoder-release-4"
def provenance_audit(known_source_groups, evaluation_groups, artifact_bytes):
overlap = sorted(known_source_groups & evaluation_groups)
return {
"known_group_overlap": overlap,
"source_groups_known": len(known_source_groups),
"target_groups_checked": len(evaluation_groups),
"artifact_sha256": hashlib.sha256(artifact_bytes).hexdigest(),
}
audit = provenance_audit(source_groups, target_test_groups, checkpoint_bytes)
assert audit["known_group_overlap"] == ["parcel-C05"]
assert len(audit["artifact_sha256"]) == 64Performance and operating cost
Intersecting S source and T target group IDs costs O(S + T) expected time and O(S + T) set memory; hashing an artifact costs O(B) time for B bytes. Manual near-duplicate review and verifying legal or consent conditions may be more expensive than either operation.
Common Mistakes
- Do not claim no source overlap when the source training manifest is unavailable.
- Do not rely on byte equality to detect all parcel-level duplicates.
- Do not transfer a checkpoint without recording its permitted use and exact version.
Read next
- Pretrained encoder and target-task contract
- Negative transfer by target slice
- Staged fine-tuning and checkpoint selection
- Transfer learning release review project
Continue the workflow: Pretraining corpus boundaries and provenance audit.
