Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Causal estimands: name the unit, treatment and missing counterfactual

Last updated: 6 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A causal effect compares outcomes under two treatment states for the same target units, although each unit reveals only one state at a time.

Specify the decision

Suppose AI Trove considers a new lesson recommendation panel. Define treatment as exposure to the new panel, the unit as learner or session, the outcome as a completed lesson within seven days and the eligible population before looking at results. An average effect for all eligible learners differs from the effect among those who clicked. Sampling and estimands keep that distinction visible.

Recognize the missing value

For one learner, we can observe completion after the new panel or after the old panel, not both at the same time. The unobserved alternative is the counterfactual. A before-and-after difference for the same learner also mixes treatment with time, fatigue or changing catalog content. Identification requires a design that supplies a credible comparison group.

Separate assignment from receipt

If a learner is assigned the new panel but never loads the page, the assigned-group effect and the effect among exposed users answer different questions. Assignment is fixed by design; exposure can depend on engagement. Intention-to-treat analysis protects the original random comparison when take-up is incomplete.

Create a decision contract

Write the population window, treatment version, unit, primary outcome, follow-up duration and exclusions. Include an interference check: a shared ranking cache might let one group alter the other’s recommendations. Test two sessions from one learner so the unit does not silently change when the same person returns.

Implementation

python
def completion_effect(treated_outcomes, control_outcomes):
    if not treated_outcomes or not control_outcomes:
        raise ValueError("both assigned groups need outcomes")
    if any(value not in (0, 1) for value in treated_outcomes + control_outcomes):
        raise ValueError("completion outcomes must be binary")
    return (sum(treated_outcomes) / len(treated_outcomes)
            - sum(control_outcomes) / len(control_outcomes))

Performance and operating cost

Computing a difference in two observed group means is O(N) time and O(1) extra space. That arithmetic is cheap; the hard work is defining the estimand and defending why the group contrast represents the missing counterfactual.

Common Mistakes

  • Do not call a before-and-after difference a treatment effect without a comparison design.
  • Do not switch from assigned users to clickers without changing the estimand.
  • Do not treat two sessions from one learner as independent units by default.

Read next

Continue the workflow: Feasible counterfactual explanations for model decisions.

Continue the workflow: Uplift targets and randomized event contracts.

ai-data
causal-inference
Storage details