Clicks and completions are observations under a serving policy, not direct measurements of every item a learner might like.
Implicit feedback: distinguish preference from what the system exposed
Read the signal carefully
A click can reflect position, title wording or current task rather than lasting preference. A completion can be missing because the learner had no time. Define positive events and their windows, then keep displayed-but-unclicked items separate from never-displayed items. The interaction contract determines which comparisons are even observable.
Account for position
An item at rank one is more likely to be noticed than an equally useful item at rank eight. Training a ranker on raw clicks can reinforce the old policy’s favorites. Log rank and policy version, evaluate by placement, and reserve controlled exploration only where the product can tolerate it. Do not call every low-ranked omission a negative preference.
Avoid duplicate events
One learner may click the same lesson repeatedly after a slow load. Deduplicate according to a defined session or impression ID, while retaining the original event audit trail. If completion arrives late, label the original impression using a fixed outcome window and version the result. Late-outcome policy keeps offline labels stable.
Inspect a counterexample
Suppose a policy displayed one lesson 47 times at rank one and another once at rank eight. Their raw click totals cannot establish which is intrinsically better. Report per-impression rates, positions and uncertainty; even those are observational unless exposure was assigned under a design that supports causal comparison.
Implementation
def unique_clicks(events):
clicked_impressions = set()
for event in events:
if event["type"] == "click":
clicked_impressions.add(event["impression_id"])
return clicked_impressionsPerformance and operating cost
Deduplicating N events with a set costs O(N) expected time and O(U) memory for U clicked impressions. A production event pipeline needs durable event IDs and retention rules; reducing data volume without preserving exposure metadata can make later evaluation invalid.
Common Mistakes
- Do not label an unseen lesson as a negative example.
- Do not interpret a raw click count without rank and impression count.
- Do not collapse repeated events before defining the impression identity.
