An observational treatment contrast requires plausible exchangeability and enough comparable treated and untreated units.
Observational overlap and weighting: know where the data can compare
Limit the claim
Suppose early adopters choose the new lesson panel while others remain on the old one. Prior activity, course stage and device may affect both selection and completion. A propensity model estimates treatment probability from pre-treatment covariates; it does not make assignment random. Unmeasured motivation can still bias the effect. Covariate roles must be argued before modeling.
Inspect common support
If every advanced learner adopts and no beginner does, the data cannot compare both treatments within those groups. Plot or summarize treatment probabilities by group and trim only under a predeclared target-population rule. Report how many units are removed and how the estimand changes. A model that outputs probabilities near zero or one is warning about weak overlap.
Bound unstable weights
Inverse-probability weighting gives large influence to rare treatment assignments. A treated learner with estimated probability 0.01 has weight 100 under a simple treatment weight. Report weight distribution and effective sample size. Truncating weights may stabilize estimates but introduces a new analysis choice; set and report the threshold before reading the effect.
Use a diagnostic fixture
Create probabilities 0.42, 0.57 and 0.01 for treated units. The last should trigger a support warning, not a quiet high-confidence estimate. Separate balance diagnostics before and after weighting. Refit no propensity feature measured after treatment and perform an unmeasured-confounding sensitivity check before drawing a causal conclusion.
Implementation
def treatment_weights(treated_flags, treatment_probabilities, minimum_probability=0.05):
if len(treated_flags) != len(treatment_probabilities) or not 0 < minimum_probability < 0.5:
raise ValueError("invalid paired inputs or support threshold")
weights = []
for treated, probability in zip(treated_flags, treatment_probabilities):
if treated not in (0, 1) or not minimum_probability <= probability <= 1 - minimum_probability:
raise ValueError("weak overlap or invalid treatment probability")
weights.append(1 / probability if treated else 1 / (1 - probability))
return weightsPerformance and operating cost
Computing N weights is O(N) time and O(N) output space. Propensity fitting can be much more expensive, but high effective weight concentration is a statistical limit that faster training cannot repair.
Common Mistakes
- Do not call propensity adjustment randomization.
- Do not hide units removed for weak overlap.
- Do not interpret a highly weighted estimate without weight and balance diagnostics.
Read next
- Confounders and covariate timing: adjust only for variables with a valid role
- Sensitivity and placebo checks: state how the causal claim could fail
- Difference-in-differences: compare changes under a defended trend assumption
- Population, estimand and sampling frame: name the quantity before calculating
Continue the workflow: Overlap and heterogeneity diagnostics for uplift models.
