Hard-negative mining selects close, reviewed nonmatches for training; the closest unverified neighbor can be a mislabeled positive or a duplicate.
Hard-negative mining without false-negative shortcuts
Why random negatives run out
Once an encoder separates different part families, random unrelated photographs satisfy the margin and contribute zero triplet loss. Nearby negatives teach finer distinctions, such as edge cuts versus surface pits on similar seals. The useful comparison is difficult because of visual resemblance, not because a missing label made it appear different. Triplet geometry makes the margin visible.
Mine inside the training boundary
Choose candidate negatives from training identities only, using an encoder checkpoint produced without validation or test identities. Refresh the mining pool at a declared cadence because the closest neighbors change during training. If an image pair was adjudicated only after a test failure, keep the test untouched and record that new label for a later cycle.
Prefer semi-hard before extreme
A semi-hard negative is farther from the anchor than the positive but still within the margin. It violates the desired separation without being the nearest possible image. The absolute closest negative can be a duplicate, mislabeled positive or unstable outlier. Inspect that tail and cap how many pairs one identity can contribute.
Audit false negatives
Review a sample of mined negatives with defect specialists. Track disagreement and the share that become positives after adjudication. Wrong negatives push equivalent repairs apart; training loss may still fall while retrieval becomes less useful. The label contract must preserve uncertain relationships.
Measure cost and coverage
A batch with B examples offers O(B squared) candidate comparisons. Full-gallery mining may require an index build and much more memory; approximate search can miss candidates. Record which identities and defect families supplied hard pairs, then measure new-identity retrieval by slice rather than trusting training loss alone.
Implementation
from math import sqrt
anchor = (0.0, 0.0)
positive = (0.30, 0.0)
reviewed_candidates = [
{"photo": "seal-83", "distance": 0.46, "relation": "negative"},
{"photo": "seal-91", "distance": 0.12, "relation": "unknown"},
{"photo": "valve-62", "distance": 1.38, "relation": "negative"},
]
def semi_hard_negatives(candidates, positive_distance, margin):
if margin <= 0:
raise ValueError("positive margin required")
return [candidate["photo"] for candidate in candidates
if candidate["relation"] == "negative"
and positive_distance < candidate["distance"] < positive_distance + margin]
selected = semi_hard_negatives(reviewed_candidates, sqrt(sum(value * value for value in positive)), 0.40)
assert selected == ["seal-83"]
assert "seal-91" not in selectedPerformance and operating cost
The filter is O(B) for B provided candidates; producing all within-batch distances is O(B squared times D) time for D dimensions and O(B squared) storage if retained. Index-based mining trades build and memory cost for faster queries. Human adjudication of close pairs is a separate cost and should be budgeted before scaling the mining process.
Common Mistakes
- Do not mine negatives from validation or test identities.
- Do not assume the nearest unverified image is a true negative.
- Do not let one frequent identity fill most hard-negative slots.
Read next
- Embedding pairs, identity labels and leakage-safe splits
- Triplet margin and normalized embedding geometry
- Embedding retrieval recall and collapse checks
- Exact versus approximate nearest-neighbor audit
Continue the workflow: Contrastive pairs, batch negatives and identity collisions.
