Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Importance-weighted risk under covariate shift

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A target-to-training density ratio can reweight historical evaluation losses toward a new input mix when the conditional outcome mechanism is stable and support overlaps.

Write the assumption before the weight

Suppose the mix of depot backlog bands changes, but the chance of a missed handoff within a band remains approximately the same. Historical evaluation losses can be weighted by target-band frequency divided by historical-band frequency. This estimates risk for the target mix only under that stability assumption and adequate overlap. The overlap guide checks the necessary support condition first.

Use evaluation rows, not training fit

The code uses already-issued predictions and mature outcomes from a held-out historical cohort. It estimates a weighted error rate using a simple band ratio supplied from separate distribution counts. It does not train a new classifier, and the tiny frequencies are illustrative. In an application, estimate ratios without leaking test outcomes, validate the conditional-stability premise and compare with newly labeled target cases when they mature. The split guide protects those roles.

Expose weight concentration

A few high-weight records can dominate the estimate. Report the largest weights and effective sample size, the squared sum of weights divided into the square of their sum. A small effective sample size means the nominal row count overstates evidence. Clip or regularize ratios only with a declared rule and report the bias this may introduce. Paired uncertainty is helpful only if the resampling unit and weights match the data design.

Know what weighting cannot repair

If a new depot changes scanner behavior, the probability of a miss at a given backlog can change. Weighting the old mix then transports the wrong conditional relationship. If a target band has zero source support, the ratio is undefined; the code raises. Reweighting also does not correct a changed label definition, late features or selective outcome recording. Concept checks require labels.

Use the estimate as one decision input

Compare weighted and unweighted risk with the same outcome maturity rule. Report interval estimates or a sensitivity range for plausible ratio error and inspect group-level harm. A weighted metric can guide where to gather evidence, but should not silently replace a future-period holdout. The response project requires that final test.

Implementation

python
# Historical validation cases: backlog band, issued alert, mature outcome.
validation_cases = [
    ("low", 0, 0), ("low", 0, 0),
    ("middle", 1, 1), ("middle", 0, 1),
    ("high", 1, 1), ("high", 1, 0),
]
source_share = {"low": 2 / 6, "middle": 2 / 6, "high": 2 / 6}
target_share = {"low": 1 / 6, "middle": 2 / 6, "high": 3 / 6}

def weighted_error(cases, historical_share, deployment_share):
    weights = []
    weighted_misses = 0.0
    for band, issued_alert, outcome in cases:
        if historical_share.get(band, 0) <= 0:
            raise ValueError("target band lacks historical support")
        weight = deployment_share[band] / historical_share[band]
        weights.append(weight)
        weighted_misses += weight * (issued_alert != outcome)
    weight_total = sum(weights)
    effective_size = weight_total ** 2 / sum(weight ** 2 for weight in weights)
    return weighted_misses / weight_total, effective_size

risk, effective_size = weighted_error(validation_cases, source_share, target_share)
assert risk == 5 / 12
assert effective_size < len(validation_cases)

Performance and operating cost

With precomputed band ratios, N weighted losses cost O(N) time and O(1) running state; the teaching code retains O(N) weights for clarity. Estimating high-dimensional density ratios is harder and unstable near weak overlap. More labeled target cases can be worth more than a complicated weighting scheme.

Common Mistakes

  • Do not describe weighted risk as unbiased when the conditional outcome relation changed.
  • Do not hide a small effective sample size behind the original row count.
  • Do not estimate ratios using target outcomes from the final evaluation period.

Read next

Continue the workflow: Offline trajectory support and simulator risk.

ai-data
machine-learning
Storage details