Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Partial identification and an assumption ledger

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Partial identification reports the range of parameter values consistent with observed data and stated assumptions when the data cannot support one point value.

Name the estimand before the interval

A queue experiment assigns branches to a new routing policy. The target is the difference in deadline-breach probability between assigned groups for every eligible ticket, including tickets whose final outcome was not captured. A responder-only difference targets a different population unless missingness is justified. Write down treatment assignment, outcome window, eligible denominator and the sign of benefit before computing any range. The estimand lesson separates the policy question from its estimator.

Distinguish three kinds of uncertainty

An identification interval reflects what the observed data and assumptions permit even with arbitrarily many similar observations. A confidence interval reflects sampling variability around a parameter or bound. A scenario range reflects selected assumptions that may not exhaust all possibilities. They are not interchangeable. The code below records a lower and upper value with its assumption label; it does not supply a statistical confidence level. Worst-case outcome bounds give a concrete construction.

Build the assumption ledger

For every tightening step, record the outcome support, missing-outcome rule, assignment mechanism, any restriction on missing-ticket risk, and the population to which it applies. A binary breach indicator has known support zero to one. That alone limits how much missing tickets can change a rate. Claiming missing tickets behave like observed tickets is much stronger and needs evidence. Keep those assumptions as separate rows rather than burying them in code.

Let wide bounds carry information

A range crossing zero may be operationally unsatisfying, but it reveals that the available outcome data cannot settle the sign under the current assumptions. Filling missing tickets with a convenient value would hide this limitation. An honest wide range can motivate targeted log recovery or a validation audit. The decision lesson shows when a narrower range would matter.

Do not promote design bounds into causal identification

If branches chose their own policy status, worst-case missing-outcome bounds on each observed group rate still do not remove selection bias. Random assignment, a defended comparison trend, or another identification design must support the treatment contrast. The code enforces ordered numeric endpoints only; the substantive ledger explains why those endpoints are valid. The attrition contrast assumes a valid assignment design.

Implementation

python
def identification_record(estimand, lower, upper, assumptions):
    if not estimand or not assumptions:
        raise ValueError("estimand and assumptions are required")
    if lower > upper:
        raise ValueError("bounds are reversed")
    return {"estimand": estimand, "lower": lower, "upper": upper,
            "assumptions": tuple(assumptions),
            "sign_identified": lower > 0 or upper < 0}

record = identification_record(
    "assigned-policy minus control breach probability",
    lower=-.14, upper=.06,
    assumptions=("binary outcome", "valid branch assignment",
                 "unobserved outcomes may be zero or one"))
assert record["lower"] <= 0 <= record["upper"]
assert record["sign_identified"] is False

Performance and operating cost

Recording one interval is O(A) time and space for A assumption strings. Computing credible bounds can be more involved; the main cost is recovering the population frame, outcome status and design evidence needed to justify the ledger.

Common Mistakes

  • Do not label an identification interval a 95% confidence interval.
  • Do not claim a causal bound from a self-selected comparison without a causal design.
  • Do not suppress a wide interval simply because it crosses the action threshold.

Read next

ai-data
data-science
Storage details