Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Hard-negative mining without false-negative shortcuts

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Hard-negative mining selects close, reviewed nonmatches for training; the closest unverified neighbor can be a mislabeled positive or a duplicate.

Why random negatives run out

Once an encoder separates different part families, random unrelated photographs satisfy the margin and contribute zero triplet loss. Nearby negatives teach finer distinctions, such as edge cuts versus surface pits on similar seals. The useful comparison is difficult because of visual resemblance, not because a missing label made it appear different. Triplet geometry makes the margin visible.

Mine inside the training boundary

Choose candidate negatives from training identities only, using an encoder checkpoint produced without validation or test identities. Refresh the mining pool at a declared cadence because the closest neighbors change during training. If an image pair was adjudicated only after a test failure, keep the test untouched and record that new label for a later cycle.

Prefer semi-hard before extreme

A semi-hard negative is farther from the anchor than the positive but still within the margin. It violates the desired separation without being the nearest possible image. The absolute closest negative can be a duplicate, mislabeled positive or unstable outlier. Inspect that tail and cap how many pairs one identity can contribute.

Audit false negatives

Review a sample of mined negatives with defect specialists. Track disagreement and the share that become positives after adjudication. Wrong negatives push equivalent repairs apart; training loss may still fall while retrieval becomes less useful. The label contract must preserve uncertain relationships.

Measure cost and coverage

A batch with B examples offers O(B squared) candidate comparisons. Full-gallery mining may require an index build and much more memory; approximate search can miss candidates. Record which identities and defect families supplied hard pairs, then measure new-identity retrieval by slice rather than trusting training loss alone.

Implementation

python
from math import sqrt

anchor = (0.0, 0.0)
positive = (0.30, 0.0)
reviewed_candidates = [
    {"photo": "seal-83", "distance": 0.46, "relation": "negative"},
    {"photo": "seal-91", "distance": 0.12, "relation": "unknown"},
    {"photo": "valve-62", "distance": 1.38, "relation": "negative"},
]

def semi_hard_negatives(candidates, positive_distance, margin):
    if margin <= 0:
        raise ValueError("positive margin required")
    return [candidate["photo"] for candidate in candidates
            if candidate["relation"] == "negative"
            and positive_distance < candidate["distance"] < positive_distance + margin]

selected = semi_hard_negatives(reviewed_candidates, sqrt(sum(value * value for value in positive)), 0.40)
assert selected == ["seal-83"]
assert "seal-91" not in selected

Performance and operating cost

The filter is O(B) for B provided candidates; producing all within-batch distances is O(B squared times D) time for D dimensions and O(B squared) storage if retained. Index-based mining trades build and memory cost for faster queries. Human adjudication of close pairs is a separate cost and should be budgeted before scaling the mining process.

Common Mistakes

  • Do not mine negatives from validation or test identities.
  • Do not assume the nearest unverified image is a true negative.
  • Do not let one frequent identity fill most hard-negative slots.

Read next

Continue the workflow: Contrastive pairs, batch negatives and identity collisions.

ai-data
machine-learning
Storage details