Build a review packet that tests whether a production breach flag changes meaning across a policy comparison and how that uncertainty changes the decision.
Project: audit a deadline-breach outcome label
Specify the measurement target
The target is whether a ticket truly breached its contracted deadline under a frozen eligibility rule. Production derives a flag from event timestamps; adjudicators reconstruct a reference label from the contract and ticket history. Record known ambiguity, clock changes and missing logs. A code value alone is not the outcome definition. The label-design lesson provides the starting point.
Draw a reviewable validation sample
Build a ticket frame by branch, period and production label. Sample within those cells using a recorded seed and inclusion probability. Review positive and negative production flags; otherwise false positives or false negatives will be unobservable. Preserve nonadjudicated cases as a separate status rather than dropping them. The sampling lesson gives a reproducible draw.
Estimate and compare error rates
Calculate sensitivity and specificity with the sampling design accounted for, and show reviewed reference-positive and reference-negative denominators. Split groups or periods if routing changes logging. If any required cell lacks reference positives or negatives, mark its rate unsupported. Report raw production rates, corrected descriptive rates and the corrected contrast without claiming the measurement correction solves treatment selection. The group-contrast lesson explains that boundary.
Stress the decision with declared assumptions
Set plausible sensitivity and specificity ranges from validation uncertainty and operational changes. Reject incompatible combinations. Compare the adjusted gap range with a policy threshold, and retain the invalid-cell count. If the sign or action changes, recommend more adjudication in the cells that most affect the result. The grid lesson shows the scenario calculation.
Deliver an evidence packet, not a green stamp
The code below checks for the artifacts needed for review; it does not compute a valid causal interval. Include the adjudication rubric, sample ledger, unresolved cases, group-specific error tables, raw and corrected rates, decision scenarios, and a separate causal-design note. If the reference process changed mid-study or a group has no validation support, stop at a scenario statement. Error-rate support is essential.
Implementation
def label_error_packet_gate(packet):
required = {"rubric", "sampling_ledger", "review_status",
"group_error_tables", "raw_rates", "corrected_rates",
"assumption_grid", "causal_design_note"}
missing = sorted(required - set(packet))
if missing:
raise ValueError("missing review artifacts: " + ", ".join(missing))
if not packet["sampling_ledger"] or not packet["group_error_tables"]:
raise ValueError("validation support is absent")
reviewed = sum(cells["reference_positive"] + cells["reference_negative"]
for cells in packet["group_error_tables"].values())
if reviewed != packet["review_status"]["complete"]:
raise ValueError("review counts disagree")
unsupported = sorted(group for group, cells in
packet["group_error_tables"].items()
if cells["reference_positive"] == 0 or
cells["reference_negative"] == 0)
return {"packet_complete": True,
"all_groups_supported": not unsupported,
"unsupported_groups": unsupported}
packet = {"rubric": "contract clock v3",
"sampling_ledger": [f"T-{ticket_number}" for ticket_number in range(73)],
"review_status": {"complete": 73, "unresolved": 0},
"group_error_tables": {"North": {"reference_positive": 13,
"reference_negative": 32},
"South": {"reference_positive": 0,
"reference_negative": 28}},
"raw_rates": {"North": .12, "South": .18},
"corrected_rates": {"North": .13},
"assumption_grid": {"valid_scenarios": 4},
"causal_design_note": "comparison trends require separate review"}
review = label_error_packet_gate(packet)
assert review == {"packet_complete": True,
"all_groups_supported": False,
"unsupported_groups": ["South"]}Performance and operating cost
The gate scans A artifact keys and G groups in O(A + G) expected time and O(G) output space. Adjudication, especially of rare reference-positive cases, is the substantial resource cost. A complete packet makes limitations visible; it does not certify the effect.
Common Mistakes
- Do not declare an unsupported group corrected because another group has validation data.
- Do not drop unresolved adjudications without reporting their pattern.
- Do not treat label correction as a replacement for causal-design review.
Read next
- Design a validation subsample for outcome labels
- Estimate outcome-label sensitivity and specificity by group
- Correct a binary outcome rate for label error
- Differential outcome-label error in group contrasts
- Misclassification assumption grids and decision ranges
Continue the workflow: Project: decide a routing policy with missing outcomes.
