Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Embedding pairs, identity labels and leakage-safe splits

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Metric learning trains a representation from similarity relationships, so the definition of a positive pair and the split unit are part of the model contract.

Decide what “same” means

A component-inspection team wants to retrieve photos of the same seal defect, not every photo of the same metal part. Two images are positive only when an adjudicated defect identity matches. Shared batch number or camera station alone is not enough; those fields may correlate with a defect and create a shortcut. Record whether the target is exact defect instance, defect family or interchangeable repair procedure before generating pairs.

Make negatives genuinely negative

A pair with different labels is not automatically a useful negative. One label may be missing, two defects may share a repair, or the same component may appear under different IDs after intake. Sample reviewed negatives and retain an “uncertain” state rather than forcing all unverified pairs apart. Weak-label conflicts are relevant when IDs are assembled from several systems.

Split by physical identity and time

All photographs of the same physical component must stay on one side of a train/test boundary. Otherwise near-duplicate views make held-out retrieval look better than new-component performance. If cameras or lighting change over time, keep a later-period test and report results by station. Group and time validation supplies the general rule.

Define the gallery at prediction time

A query image is compared with a gallery of known examples. Store gallery revision, eligibility rules, encoder version and feature extraction time. Test queries must not have their own duplicate in the gallery unless duplicate recognition is the stated product task. Feature availability remains a hard boundary.

Count independent identities

Ten thousand nearly identical frames from one faulty part are not ten thousand independent positive examples. Report identities, source batches, positive pairs and uncertain labels separately. A per-image train/test split can produce a large metric that has little meaning for unseen equipment. The release review carries these counts.

Implementation

python
inspection_photos = [
    {"photo": "seal-47-front", "component": "seal-47", "defect": "cut-edge", "captured": 912},
    {"photo": "seal-47-side", "component": "seal-47", "defect": "cut-edge", "captured": 915},
    {"photo": "seal-83-front", "component": "seal-83", "defect": "surface-pit", "captured": 926},
    {"photo": "seal-91-front", "component": "seal-91", "defect": None, "captured": 937},
]

def pair_relation(first, second):
    if first["defect"] is None or second["defect"] is None:
        return "unknown"
    if first["component"] == second["component"]:
        return "same-component"  # Exclude from unseen-component evaluation.
    return "positive" if first["defect"] == second["defect"] else "negative"

assert pair_relation(inspection_photos[0], inspection_photos[1]) == "same-component"
assert pair_relation(inspection_photos[0], inspection_photos[2]) == "negative"
assert pair_relation(inspection_photos[0], inspection_photos[3]) == "unknown"

Performance and operating cost

Pair classification is O(1) per candidate pair, but enumerating every pair across N photos costs O(N squared) time and storage if materialized. Keep identity metadata and sample pairs within training folds. Adjudicating ambiguous defects is often more expensive than computing distances, yet it determines whether the geometry means anything.

Common Mistakes

  • Do not define similarity by a field that is merely easy to collect.
  • Do not place views of one physical component in both training and test.
  • Do not turn missing defect labels into verified negatives.

Read next

ai-data
machine-learning
Storage details