A release packet connects view validity, corpus lineage, frozen transfer tests and operating cost before an unlabeled-image encoder is accepted for damage review.
Project: release review for an unlabeled parcel-image encoder
Freeze the evaluation claim
The decision is whether a parcel image should enter manual damage review at intake. Reserve later parcels and new-depot slices before gathering the pretraining pool. Record the amount of labeled data allowed for each candidate. The release claim concerns future parcels, so parcel identity must not cross the pretraining and final-test boundary. Corpus provenance provides the first gate.
Build two pretraining candidates
For a contrastive candidate, review paired crops to ensure tears and crushed corners remain visible, then filter known same-parcel negatives. For a masked-reconstruction candidate, choose a mask policy that cannot be solved through copied metadata and measure its defect-overlap rate. Compare both with a no-pretraining encoder. Pair audit and mask design define the failure modes.
Run a fixed downstream protocol
Freeze each encoder, fit a linear head on the same labeled training parcels and select thresholds on the same development set. Report per-depot recall, false review volume and confidence intervals or repeated-seed spread with support counts. Fine-tune separately if resources allow, but keep that result distinct from the frozen probe. Probe evidence tests transfer.
Measure serving and rollback
Benchmark image decoding, preprocessing and encoder inference together on the target device. Keep the existing damage-review rule available during a shadow phase, record drift in camera mix and false-review load, and state who can disable the new encoder. A smaller pretext loss is not a rollback criterion; downstream slice quality and operating cost are.
Make a conditional decision
The sample packet holds release because one test parcel overlaps the pretraining corpus, despite acceptable latency and overall recall. Exclude the parcel, rerun the frozen protocol on a clean untouched set and review rare-defect slices before reconsidering. Do not erase the failed audit from the record.
Implementation
release_packet = {
"unresolved_overlap_ids": ["dock-83"],
"view_damage_audit_passed": True,
"rare_defect_recall": 0.81,
"minimum_rare_defect_recall": 0.79,
"p95_full_path_ms": 58,
"maximum_p95_ms": 72,
"rollback_owner_assigned": True,
}
def review_pretraining_release(packet):
blockers = []
if packet["unresolved_overlap_ids"]:
blockers.append("pretraining and final-test identity overlap")
if not packet["view_damage_audit_passed"]:
blockers.append("view audit failed")
if packet["rare_defect_recall"] < packet["minimum_rare_defect_recall"]:
blockers.append("rare-defect recall below minimum")
if packet["p95_full_path_ms"] > packet["maximum_p95_ms"]:
blockers.append("serving budget exceeded")
if not packet["rollback_owner_assigned"]:
blockers.append("rollback owner missing")
return {"release": not blockers, "blockers": blockers}
decision = review_pretraining_release(release_packet)
assert decision["release"] is False
assert decision["blockers"] == ["pretraining and final-test identity overlap"]Performance and operating cost
The packet gate is O(K) in the number of recorded overlap IDs and O(1) for its remaining scalar checks. The release evaluation itself requires multiple encoder training runs, fixed-label probes, identity audits and target-device benchmarks. Set a compute ceiling before comparing candidates.
Common Mistakes
- Do not approve a model because its pretext metric improved.
- Do not reuse a contaminated final test after choosing a model from it.
- Do not omit identity, rare-defect and serving-cost gates from the decision.
