A causal effect compares outcomes under two treatment states for the same target units, although each unit reveals only one state at a time.
Causal estimands: name the unit, treatment and missing counterfactual
Specify the decision
Suppose AI Trove considers a new lesson recommendation panel. Define treatment as exposure to the new panel, the unit as learner or session, the outcome as a completed lesson within seven days and the eligible population before looking at results. An average effect for all eligible learners differs from the effect among those who clicked. Sampling and estimands keep that distinction visible.
Recognize the missing value
For one learner, we can observe completion after the new panel or after the old panel, not both at the same time. The unobserved alternative is the counterfactual. A before-and-after difference for the same learner also mixes treatment with time, fatigue or changing catalog content. Identification requires a design that supplies a credible comparison group.
Separate assignment from receipt
If a learner is assigned the new panel but never loads the page, the assigned-group effect and the effect among exposed users answer different questions. Assignment is fixed by design; exposure can depend on engagement. Intention-to-treat analysis protects the original random comparison when take-up is incomplete.
Create a decision contract
Write the population window, treatment version, unit, primary outcome, follow-up duration and exclusions. Include an interference check: a shared ranking cache might let one group alter the other’s recommendations. Test two sessions from one learner so the unit does not silently change when the same person returns.
Implementation
def completion_effect(treated_outcomes, control_outcomes):
if not treated_outcomes or not control_outcomes:
raise ValueError("both assigned groups need outcomes")
if any(value not in (0, 1) for value in treated_outcomes + control_outcomes):
raise ValueError("completion outcomes must be binary")
return (sum(treated_outcomes) / len(treated_outcomes)
- sum(control_outcomes) / len(control_outcomes))Performance and operating cost
Computing a difference in two observed group means is O(N) time and O(1) extra space. That arithmetic is cheap; the hard work is defining the estimand and defending why the group contrast represents the missing counterfactual.
Common Mistakes
- Do not call a before-and-after difference a treatment effect without a comparison design.
- Do not switch from assigned users to clickers without changing the estimand.
- Do not treat two sessions from one learner as independent units by default.
Read next
- Confounders and covariate timing: adjust only for variables with a valid role
- Randomized assignment: check balance and retain every assigned unit
- Population, estimand and sampling frame: name the quantity before calculating
- Experiment design: assign the right unit and guard against interference
Continue the workflow: Feasible counterfactual explanations for model decisions.
Continue the workflow: Uplift targets and randomized event contracts.
