A sensitivity check changes uncertain assumptions while keeping the target question and observed evidence fixed.
Sensitivity analysis: find which assumptions can reverse a decision
List uncertainties before scenarios
A support-cohort estimate may depend on unresolved cases, imperfect closure labels and an incomplete self-service channel. Name each uncertainty and its plausible range using audit evidence or a documented operating limit. Do not select the range after seeing which result favors a proposed launch. Proxy error and missing outcomes are separate dimensions.
Calculate a decision boundary
Suppose 86 of 143 fully observed eligible cases met a two-day service target, while 19 cases lack final confirmation. If all 19 failed, the rate over 162 cases is 86/162, about 53.1%. If all 19 succeeded, it is 105/162, about 64.8%. A policy requires at least 60%. That decision can flip under the missing outcomes; the correct report is a range and the number of additional confirmations needed to narrow it.
Change one assumption, then combine
Vary follow-up classification, proxy-label correction and frame coverage individually first. This reveals which uncertainty drives the conclusion. Then evaluate coherent combinations: difficult cases may both remain unresolved and be absent from the easy-to-query export, so independent best-case assumptions are implausible. Keep each scenario’s denominator and cohort rule visible. If weights change, recompute the weighted rate rather than editing a rounded headline.
Separate uncertainty types
A bootstrap interval describes sampling variation under its design assumptions. A sensitivity range describes how a result moves under stated deviations from uncertain assumptions. Neither substitutes for the other. An interval can be narrow around a biased proxy, and a wide sensitivity range can persist with thousands of observations. Resampling at the right unit addresses only one part of the problem.
Make the next measurement actionable
If the decision remains the same across credible scenarios, record that stability and the assumptions tested. If it flips, state the smallest new observation that would resolve the uncertainty: recontact a random subset of the 19, review closure labels by channel, or recover missing intake IDs. Order that work by expected decision value, not by how easy a chart is to draw.
Implementation
def binary_missing_outcome_range(success_count, observed_count, missing_count):
if not 0 <= success_count <= observed_count or missing_count < 0:
raise ValueError("invalid cohort counts")
total = observed_count + missing_count
if total == 0:
raise ValueError("empty cohort")
return success_count / total, (success_count + missing_count) / total
low, high = binary_missing_outcome_range(86, 143, 19)
assert round(low, 3) == 0.531 and round(high, 3) == 0.648Performance and operating cost
A single binary bound takes O(1) time and space. A grid of A × B × C assumption settings costs O(ABC) evaluations and can grow quickly; predeclare a small, interpretable range and preserve the scenario inputs with each output.
Common Mistakes
- Do not describe an assumption range as a confidence interval.
- Do not choose only scenarios that support the desired decision.
- Do not vary a numerator without updating the corresponding denominator and cohort definition.
Read next
- Proxy measurements and label error in operational datasets
- Right censoring and time-to-event analysis for open cases
- Bootstrap intervals: estimate uncertainty at the right sampling unit
- Project: audit support resolution with open cases and imperfect labels
Continue the workflow: Derived metrics: carry input bounds and shared error into the result.
Continue the workflow: Constraint feasibility and slack: reject impossible plans early.
Continue the workflow: Missing-not-at-random sensitivity: test what unseen outcomes could change.
Continue the workflow: Misclassification assumption grids and decision ranges.
Continue the workflow: Bounded missing-outcome risk scenarios.
