A release gate must account for measured group harm, incomplete outcomes and the cost of an operational change.
Cohort release gates: compare harm, coverage and uncertainty together
Select a decision rule before comparison
A candidate claim model changes its score threshold. Define the target harm, overall capacity, group minimum counts and maximum acceptable group gap before seeing its holdout result. A strict equality target is not automatically the right metric: the review team cares about missed genuine claims while also being able to handle the alert volume. Record who chose the policy and what population it covers. Denominator contracts make the comparison interpretable.
Separate lack of evidence from a pass
If a region has only a few mature genuine claims, the apparent gap is unstable. Hold the cohort judgment as insufficient evidence and use an approved audit sample or longer observation window. Do not bury that region in an overall metric. If labels are collected mostly for flagged claims, estimate the observation bias before calling an unflagged case negative. Audit sampling can reveal the unseen part of the workflow.
Test a candidate across constraints
Compare candidate and incumbent with the same label definition, review-capacity limit, cohort membership and time window. Check overall recall, each region’s false-negative rate, alert volume, missing-region share and confidence or count threshold. A candidate may lower the maximum group gap by making every region worse. Reject that outcome even when the parity number improves. Slice gates prevent an aggregate win from masking local harm.
Release with remeasurement
If the evidence passes, canary by stable assignment and keep model revision in the event ledger. Watch group counts and mature outcomes after the canary; delayed labels mean early dashboards cannot claim a quality win. Set an owner, a data-collection action for insufficient cohorts and a rollback trigger for a confirmed harm increase. Delayed-label monitoring defines when the quality window closes. The project holds a seemingly better candidate whose unknown-region share grows.
Implementation
def cohort_release_gate(candidate, limits):
if candidate["unknown_region_share"] > limits["maximum_unknown_share"]:
return "hold:coverage"
if min(candidate["mature_positives"].values()) < limits["minimum_positives"]:
return "hold:insufficient-evidence"
if candidate["overall_recall"] < limits["minimum_recall"]:
return "hold:overall-harm"
rates = list(candidate["false_negative_rates"].values())
if max(rates) - min(rates) > limits["maximum_gap"]:
return "review:cohort-gap"
if candidate["daily_alerts"] > limits["alert_capacity"]:
return "hold:capacity"
return "canary:remeasure"
limits = {"maximum_unknown_share": 0.08, "minimum_positives": 40,
"minimum_recall": 0.84, "maximum_gap": 0.12, "alert_capacity": 420}
candidate = {"unknown_region_share": 0.11,
"mature_positives": {"north": 68, "south": 57},
"overall_recall": 0.88,
"false_negative_rates": {"north": 0.10, "south": 0.16},
"daily_alerts": 390}
assert cohort_release_gate(candidate, limits) == "hold:coverage"
assert cohort_release_gate({**candidate, "unknown_region_share": 0.04},
limits) == "canary:remeasure"
Performance and operating cost
The gate scans g cohorts in O(g) time and O(g) temporary space. Longer label windows, approved audit sampling and human review cost far more than the metric calculation. A count threshold is a triage rule, not a substitute for an uncertainty estimate or an investigation of missing outcomes.
Common Mistakes
- Optimizing only the disparity number while every group loses recall.
- Passing a tiny cohort because its observed rate happens to look good.
- Treating missing group membership as random without checking its pattern.
- Declaring the canary successful before delayed labels mature.
Read next
- Cohort outcome monitoring: keep denominators and observation windows visible
- Project: review regional claim outcomes before a model release
- Slice quality gates when labels are sparse or delayed
- Model monitoring: separate input drift, data faults and delayed outcomes
- Selective labels: measure what the model never lets reviewers see
