A target-to-training density ratio can reweight historical evaluation losses toward a new input mix when the conditional outcome mechanism is stable and support overlaps.
Importance-weighted risk under covariate shift
Write the assumption before the weight
Suppose the mix of depot backlog bands changes, but the chance of a missed handoff within a band remains approximately the same. Historical evaluation losses can be weighted by target-band frequency divided by historical-band frequency. This estimates risk for the target mix only under that stability assumption and adequate overlap. The overlap guide checks the necessary support condition first.
Use evaluation rows, not training fit
The code uses already-issued predictions and mature outcomes from a held-out historical cohort. It estimates a weighted error rate using a simple band ratio supplied from separate distribution counts. It does not train a new classifier, and the tiny frequencies are illustrative. In an application, estimate ratios without leaking test outcomes, validate the conditional-stability premise and compare with newly labeled target cases when they mature. The split guide protects those roles.
Expose weight concentration
A few high-weight records can dominate the estimate. Report the largest weights and effective sample size, the squared sum of weights divided into the square of their sum. A small effective sample size means the nominal row count overstates evidence. Clip or regularize ratios only with a declared rule and report the bias this may introduce. Paired uncertainty is helpful only if the resampling unit and weights match the data design.
Know what weighting cannot repair
If a new depot changes scanner behavior, the probability of a miss at a given backlog can change. Weighting the old mix then transports the wrong conditional relationship. If a target band has zero source support, the ratio is undefined; the code raises. Reweighting also does not correct a changed label definition, late features or selective outcome recording. Concept checks require labels.
Use the estimate as one decision input
Compare weighted and unweighted risk with the same outcome maturity rule. Report interval estimates or a sensitivity range for plausible ratio error and inspect group-level harm. A weighted metric can guide where to gather evidence, but should not silently replace a future-period holdout. The response project requires that final test.
Implementation
# Historical validation cases: backlog band, issued alert, mature outcome.
validation_cases = [
("low", 0, 0), ("low", 0, 0),
("middle", 1, 1), ("middle", 0, 1),
("high", 1, 1), ("high", 1, 0),
]
source_share = {"low": 2 / 6, "middle": 2 / 6, "high": 2 / 6}
target_share = {"low": 1 / 6, "middle": 2 / 6, "high": 3 / 6}
def weighted_error(cases, historical_share, deployment_share):
weights = []
weighted_misses = 0.0
for band, issued_alert, outcome in cases:
if historical_share.get(band, 0) <= 0:
raise ValueError("target band lacks historical support")
weight = deployment_share[band] / historical_share[band]
weights.append(weight)
weighted_misses += weight * (issued_alert != outcome)
weight_total = sum(weights)
effective_size = weight_total ** 2 / sum(weight ** 2 for weight in weights)
return weighted_misses / weight_total, effective_size
risk, effective_size = weighted_error(validation_cases, source_share, target_share)
assert risk == 5 / 12
assert effective_size < len(validation_cases)Performance and operating cost
With precomputed band ratios, N weighted losses cost O(N) time and O(1) running state; the teaching code retains O(N) weights for clarity. Estimating high-dimensional density ratios is harder and unstable near weak overlap. More labeled target cases can be worth more than a complicated weighting scheme.
Common Mistakes
- Do not describe weighted risk as unbiased when the conditional outcome relation changed.
- Do not hide a small effective sample size behind the original row count.
- Do not estimate ratios using target outcomes from the final evaluation period.
Read next
- Covariate shift and support overlap
- Concept change with mature outcomes
- Group and time validation: split by the failure you expect in production
- Distribution shift response project
Continue the workflow: Offline trajectory support and simulator risk.
