Pairwise learning compares differently graded candidates within the same query group; equal-grade pairs and cross-query pairs do not express the desired ordering.
Pairwise ranking loss, ties and useful comparisons
Train on within-request preferences
When an approved seal-repair page has grade 3 and a parts index has grade 1 for the same request, the ranker should assign the repair page a larger score. Scores need not be calibrated probabilities. The code computes logistic loss on a small, declared set of preferred pairs. It deliberately skips tied grades and never pairs pages from different requests.
Know what the objective optimizes
A pairwise logistic objective penalizes reversed and uncertain comparisons, but it does not directly optimize the whole displayed list. Its gradient can improve a pair below the fold that no technician sees. A list-aware method may weight swaps near the top more heavily. NDCG at a cutoff measures the result at the decision surface.
Control pair explosion
A request with M candidates can yield O(M squared) ordered comparisons. Sample within the query, retain positive and negative grade coverage, and record the sampling rule. Long requests should not dominate training merely because they generate more pairs; normalize or weight at query level. The number of effective pairs matters more than raw row count.
Handle ties and contradictory labels
Two pages with the same adjudicated grade are not ordered by that rubric. If reviewers disagree, retain the adjudication history or a distribution rather than inventing a preference. Repeated near-identical pages may need deduplication before pair construction, because otherwise a training set can overrepresent one procedure revision.
Keep the baseline honest
Compare against a stable heuristic such as exact equipment match followed by document freshness. Optimize model settings only on development requests, then evaluate once on untouched later requests. Untouched testing protects the final comparison; the release project joins it with serving cost.
Implementation
from math import exp, log1p
search_candidates = [
{"request": "pump-47", "document": "seal-a7", "grade": 3, "score": 1.9},
{"request": "pump-47", "document": "parts-b4", "grade": 1, "score": 0.4},
{"request": "pump-47", "document": "check-c2", "grade": 1, "score": 0.7},
{"request": "valve-62", "document": "isolation-c9", "grade": 2, "score": 0.3},
]
def preferred_pair_losses(candidates):
groups = {}
for candidate in candidates:
groups.setdefault(candidate["request"], []).append(candidate)
losses = []
for group in groups.values():
for first_index, first in enumerate(group):
for second in group[first_index + 1:]:
if first["grade"] == second["grade"]:
continue
better, worse = (first, second) if first["grade"] > second["grade"] else (second, first)
margin = better["score"] - worse["score"]
losses.append(log1p(exp(-margin)))
return losses
losses = preferred_pair_losses(search_candidates)
assert len(losses) == 2
assert all(0 < loss < log1p(1) for loss in losses)Performance and operating cost
Grouping takes O(N) expected time and O(N) memory. Enumerating every pair costs O(sum of group-size squared) time and O(P) memory if all P losses are retained; stream or sample pairs for large requests. Training costs depend on the chosen scorer, and pair count does not equal the number of independent search events.
Common Mistakes
- Do not form a preference from equal grades.
- Do not compare scores from unrelated query groups as if the labels shared one scale.
- Do not let a handful of long queries determine the objective by pair count alone.
