Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit runbook retrieval pairs before training

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build a versioned query-passage training manifest, review difficult negatives and hold an export when an answer-bearing family is mislabeled.

Specify the retrieval job

A runbook retriever must find the current gateway-west rollback instruction for several question forms. Collect reviewed query-passage positives with exact source spans and revisions. Mine confusing gateway-east, obsolete and non-answer sections as candidate negatives. Do not train from the final answer alone; the passage judgment is the unit the retriever learns. Positive-pair contracts keep the label tied to a source revision.

Build a review fixture

Include paraphrased questions, code-switched queries, copied runbook sections and a changed numeric threshold. Give reviewers the original passage, service and environment without displaying the model’s proposed label. Record positive, negative or unresolved, plus a reason. Family IDs connect copies and revisions; a negative for one query may be positive for another. Split by incident and source family before comparing model changes. Family-aware splits prevent an almost identical passage from leaking across evaluation partitions.

Enforce manifest invariants

Check that no query-family pair is both positive and negative in the same source revision. Reject empty positive sets and unresolved candidates presented as labeled negatives. Enforce access policy before exporting text. The example gate catches direct contradictions; human review is still required to find a relevant passage the judgment set missed. Negative review keeps high-ranked unknowns in a queue until a decision is supported.

Report a bounded result

Evaluate candidate recall, first relevant rank and hard-negative errors on a frozen, independently judged set. Report false-negative corrections and reviewer disagreement. Hold training if the label manifest contradicts itself or contains restricted text. After repair, train and re-evaluate without tuning against the held-out judgments. The released artifact includes manifest version, corpus snapshot, index version and known coverage limits.

Implementation

python
def retrieval_manifest_conflicts(labels):
    judgments = {}
    for row in labels:
        key = (row["query_id"], row["family_id"], row["revision_id"])
        if row["label"] not in {"positive", "negative", "unresolved"}:
            raise ValueError("unknown retrieval label")
        judgments.setdefault(key, set()).add(row["label"])
    return {key: sorted(values) for key, values in judgments.items()
            if "positive" in values and "negative" in values}

labels = [
    {"query_id": "rollback-47", "family_id": "runbook-west",
     "revision_id": "r8", "label": "positive"},
    {"query_id": "rollback-47", "family_id": "runbook-west",
     "revision_id": "r8", "label": "negative"},
]
assert retrieval_manifest_conflicts(labels) == {
    ("rollback-47", "runbook-west", "r8"): ["negative", "positive"]}
labels[1]["label"] = "unresolved"
assert retrieval_manifest_conflicts(labels) == {}

Performance and operating cost

Checking n label rows takes expected O(n) time and O(n) space for grouped judgments; sorting the small conflicting label sets adds bounded work here. A contradiction-free manifest may still contain false negatives, missing positives or access violations, so this is one release check rather than a complete relevance audit.

Common Mistakes

  • Exporting unresolved candidates as negatives.
  • Ignoring contradictory labels for the same query-family revision.
  • Evaluating with copied runbook sections in both training and test.
  • Claiming the manifest is clean only because its explicit conflicts were removed.

Read next

Continue the workflow: Multi-step questions: decompose answer slots and dependencies.

ai-data
natural-language-processing
Storage details