Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Retrieval negatives: hard candidates and false-negative risk

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A high-ranked unjudged passage is a useful review candidate, not automatically a negative training label.

Start from what is known

A hard negative looks relevant to the retriever yet does not satisfy the query. The gateway-east rollback section may resemble a gateway-west answer while referring to the wrong service. That is a defensible negative after review. An unjudged gateway-west passage might answer the query in different words; forcing it into the negative class trains the retriever to suppress useful evidence. Keep judged negatives, unjudged candidates and known positives as separate states. Positive-pair identity anchors the known answer-bearing passages.

Review candidate origin

Mine candidates from the current retriever, lexical search and nearby sections of the same runbook. Record score, source revision, access level, language and why the passage was nominated. Exclude exact copies and known positive families before labeling; route semantic near-matches to a reviewer. A candidate from a later runbook revision can be a valid future positive rather than a safe negative for an old query. Text-family review catches copied or revised answer spans.

Guard the optimization signal

Compare query and candidate under the same service, environment, version and requested action. Do not equate absence from a sparse judgment set with non-relevance. A false negative is particularly harmful when it is a strong answer-bearing passage, because training pushes it away from the query. Keep a small pool of randomly sampled negatives for coverage and review the difficult ones more carefully. Label uncertainty explicitly rather than pretending every mined candidate has a binary truth.

Evaluate after mining

Freeze an independent judgment set before tuning the retriever. Measure recall of all judged positives, errors on hard negatives and the fraction of mined candidates later found relevant. Slice by language pair, service and source revision. A rising training score can coexist with worse retrieval if false negatives accumulate. The label audit project blocks a training export when an answer-bearing passage appears in both positive and negative roles.

Implementation

python
def negative_review_queue(query_id, candidates, positive_family_ids,
                          allowed_access):
    queue = []
    for candidate in candidates:
        if candidate["access_level"] not in allowed_access:
            continue
        if candidate["family_id"] in positive_family_ids:
            continue
        if candidate["judgment"] == "positive":
            continue
        queue.append({"query_id": query_id,
                      "passage_id": candidate["passage_id"],
                      "state": "review-needed"})
    return queue

candidates = [
    {"passage_id": "runbook-82", "family_id": "rollback-west",
     "access_level": "operations", "judgment": "unjudged"},
    {"passage_id": "runbook-91", "family_id": "rollback-east",
     "access_level": "operations", "judgment": "unjudged"},
]
queue = negative_review_queue("rollback-47", candidates,
                              {"rollback-west"}, {"operations"})
assert queue == [{"query_id": "rollback-47", "passage_id": "runbook-91",
                  "state": "review-needed"}]

Performance and operating cost

Scanning c candidates takes expected O(c) time with O(c) output space. The function only proposes review candidates; it does not issue negative labels. Candidate mining through an index, source-family comparison and human adjudication carry separate costs that rise with corpus size and language coverage.

Common Mistakes

  • Treating every unjudged search result as a negative.
  • Pushing a copied positive passage away because it has a different URL.
  • Mining from a newer revision while ignoring changed answer truth.
  • Letting access-restricted candidates leak into a training export.

Read next

ai-data
natural-language-processing
Storage details