A low pretraining loss is insufficient; feature spread, simple downstream probes and target-slice results show whether the encoder retained information worth transferring.
Representation collapse and frozen linear probes
Inspect variance before training a head
An encoder that maps every parcel to nearly the same vector has little useful capacity, even if a training diagnostic looks stable. Check per-dimension spread on held-out unlabeled data and inspect duplicate or near-duplicate rates. This is only a diagnostic: high variance can reflect camera or depot identity rather than damage.
Freeze the encoder for a probe
Fit a small linear classifier on labeled training embeddings while keeping encoder weights fixed. Tune its regularization and threshold on development data, then evaluate once on later held-out parcels. The frozen probe tests separability with little adaptation. It is not a full replacement for fine-tuning, which may recover information that a linear head cannot use. Probe setup expands the comparison.
Compare to simple baselines
Use a no-pretraining encoder of the same size and a conventional feature baseline where practical. Keep label budgets identical. If the pretrained encoder wins only after seeing more labeled examples, it does not establish a label-efficiency gain. Plot performance against labeled sample count and include variation across seeds. Learning curves show the pattern.
Test the actual slices
Report damage recall, false review rate and calibration by depot, package surface and rare defect. A global probe score can improve while the damaged-corner cohort regresses. Compare the same later-period parcel IDs for every candidate and preserve support counts.
Distinguish a diagnostic from a release gate
The simple spread check below catches exact coordinate collapse in a toy set. It cannot prove downstream utility or fairness. Release needs a paired downstream test, provenance audit and measured serving cost. The project combines those gates.
Implementation
embeddings = {
"dock-47": (0.9, -0.2, 0.4),
"dock-62": (-0.7, 0.5, 0.1),
"dock-83": (0.2, 0.1, -0.6),
}
def coordinate_spread(vectors):
points = list(vectors.values())
if not points or len({len(point) for point in points}) != 1:
raise ValueError("nonempty, equal-width vectors required")
return tuple(max(column) - min(column) for column in zip(*points))
spread = coordinate_spread(embeddings)
assert spread == (1.6, 0.7, 1.0)
assert coordinate_spread({"one": (0.2, 0.2), "two": (0.2, 0.2)}) == (0.0, 0.0)Performance and operating cost
Computing coordinate ranges for N vectors of width D costs O(ND) time and O(D) auxiliary memory if streamed. Producing embeddings costs encoder inference over the audit set; a probe adds training cost. Exact collapse is cheap to detect, but meaningful transfer requires labeled holdouts and slice review.
Common Mistakes
- Do not use feature variance as a substitute for downstream quality.
- Do not tune the probe on the final test.
- Do not compare candidates trained with unequal label budgets.
