Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Click position bias and support for ranking evaluation

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Search clicks are affected by result position and exposure, so they cannot be treated as direct relevance labels without an observation model and adequate support.

Log what could be seen

A technician can click only a displayed result, and a result at the top is more likely to be examined. Keep the presented order, candidate IDs, snippet version, device, session and displayed timestamp. A click event without its display context cannot distinguish relevance from visibility. Exposure logging defines a related product-event contract.

Separate the estimands

Editorial relevance asks whether a manual page answers a request. Click probability asks whether someone selects it under a particular interface and rank. Successful repair may require a later verified outcome. These signals should be reported separately; otherwise a model can optimize for attention-grabbing snippets while sending technicians to less useful instructions.

Randomization needs guardrails

Small controlled swaps among safe, eligible candidates can estimate position effects. Do not randomize safety-critical procedures into an unreviewed first position. Log the assignment probability and the exact eligible set for every randomized display. If a target policy chooses a page in contexts never exposed by the logging policy, counterfactual evaluation cannot identify its outcome from those logs. Bandit support spells out the condition.

Inspect weights before trusting a correction

Inverse-propensity weighting divides an observed outcome by its display probability. Tiny probabilities produce high-variance estimates; clipping changes the estimand and must be disclosed. Report the effective sample size and a sensitivity range, not just a corrected mean. The code audits weights from a made-up randomized logging packet; it does not prove unbiased relevance.

Retain a human-judged test set

A fixed, independently judged query set can reveal regressions caused by click-derived labels and changing UI placement. Refresh it when the document corpus changes, and keep the old set for a stability trend. Per-query NDCG and release review use that set.

Implementation

python
display_events = [
    {"request": "pump-47", "rank": 1, "clicked": 1, "display_probability": 0.50},
    {"request": "valve-62", "rank": 2, "clicked": 0, "display_probability": 0.25},
    {"request": "motor-83", "rank": 3, "clicked": 1, "display_probability": 0.20},
]

def weight_diagnostics(events):
    if not events or any(not 0 < event["display_probability"] <= 1 for event in events):
        raise ValueError("positive logged probabilities required")
    weights = [1 / event["display_probability"] for event in events]
    effective_count = sum(weights) ** 2 / sum(weight * weight for weight in weights)
    weighted_click_rate = sum(weight * event["clicked"] for weight, event in zip(weights, events)) / sum(weights)
    return effective_count, weighted_click_rate, max(weights)

effective_count, corrected_rate, largest_weight = weight_diagnostics(display_events)
assert round(effective_count, 3) == 2.689
assert round(corrected_rate, 3) == 0.636
assert largest_weight == 5.0

Performance and operating cost

The weight audit uses O(N) time and O(N) temporary memory as written; streaming sums reduces memory to O(1). Randomized logging consumes user traffic and operational review capacity. Weighting only addresses the modeled exposure process; unlogged eligibility, snippet changes and missing outcomes can still bias the estimate.

Common Mistakes

  • Do not label every unclicked or unshown page irrelevant.
  • Do not use inverse weights when display probabilities were not logged or are zero.
  • Do not treat a high corrected click rate as proof of safe task completion.

Read next

ai-data
machine-learning
Storage details