An assumption grid recalculates an outcome contrast across a prespecified set of plausible sensitivity and specificity values and reports where the decision changes.
Misclassification assumption grids and decision ranges
Make uncertainty in error rates visible
A small validation audit rarely pins down sensitivity and specificity to one exact value. Build plausible ranges from the audit and operational knowledge, separately for groups whose logging differs. Freeze the grid before examining which cell best supports a favored policy. Each cell is a scenario, not a posterior draw or a confidence bound. The validation lesson explains where the central rates come from.
Reject impossible cells explicitly
For every sensitivity-specificity pair, invert the observed rate. If the sum is at most one or the implied true rate lies outside zero to one, mark the cell invalid; do not clip it and include it in the range. The code loops over declared assumptions for two groups and retains only compatible combinations. The resulting minimum and maximum adjusted gaps describe the evaluated assumption set only. They are not an identified interval unless that set is defensibly exhaustive.
Connect the range to an action threshold
Suppose a queue rollout requires at least a three-point reduction in true breach rate to justify cost. Count scenarios where the adjusted treated-minus-comparison gap is at most -0.03. If the action changes across plausible assumptions, prioritize a better validation audit or a narrower rollout. Do not average scenario outputs without explicit weights for the assumptions. Decision stability covers the broader principle.
Do not omit sampling uncertainty
This grid varies assumed accuracy but holds observed rates fixed. In a formal probabilistic bias analysis, validation and production sampling uncertainty would be propagated too, with care for dependence from an internal validation sample. A coarse grid can miss a decision boundary between selected values. Increase resolution around the boundary, but disclose the revised grid and reason. Group-specific correction is the calculation inside each cell.
Keep the failure stories separate
A result stable to label error can still fail because of confounding or a contaminated comparison group. Conversely, weak label validation can dominate even an otherwise careful study. Show the raw contrast, central corrected contrast, scenario range, invalid-cell count and action threshold together. The audit project turns these into a review packet.
Implementation
from itertools import product
def corrected_rate(observed, sensitivity, specificity):
if not all(0 <= value <= 1 for value in
(observed, sensitivity, specificity)):
return None
information = sensitivity + specificity - 1
if information <= 0:
return None
rate = (observed + specificity - 1) / information
return rate if 0 <= rate <= 1 else None
def contrast_assumption_grid(treated_observed, comparison_observed,
treated_options, comparison_options, action_gap):
gaps = []
invalid = 0
for treated_error, comparison_error in product(treated_options, comparison_options):
treated_rate = corrected_rate(treated_observed, *treated_error)
comparison_rate = corrected_rate(comparison_observed, *comparison_error)
if treated_rate is None or comparison_rate is None:
invalid += 1
continue
gaps.append(treated_rate - comparison_rate)
if not gaps:
raise ValueError("no compatible assumptions")
return {"range": (min(gaps), max(gaps)),
"action_scenarios": sum(gap <= action_gap for gap in gaps),
"valid_scenarios": len(gaps), "invalid_scenarios": invalid}
treated_errors = [(.82, .97), (.88, .99)]
comparison_errors = [(.91, .94), (.96, .97)]
audit = contrast_assumption_grid(.12, .18, treated_errors,
comparison_errors, action_gap=-.03)
assert audit["valid_scenarios"] == 4
assert audit["invalid_scenarios"] == 0
assert audit["range"][0] <= audit["range"][1]Performance and operating cost
With T treated and C comparison assumption pairs, direct evaluation costs O(T × C) time and O(T × C) space for retained gaps. A formal uncertainty model has additional sampling and simulation cost; grid size alone does not establish validity.
Common Mistakes
- Do not label a handpicked scenario range a confidence interval.
- Do not include clipped or otherwise infeasible corrected rates.
- Do not average scenarios without declared weights or probabilities.
