A group audit computes outcomes and error rates separately for relevant populations, then investigates differences in labels, exposure and intervention cost before changing policy.
Group error gaps and policy audit
Specify the groups and action
A handoff alert sends a case to an overloaded review desk. Compare false-negative and false-positive rates by site, shift or another justified group; if people are affected, relevant protected-group analysis may also be required under the applicable governance process. A selection-rate difference alone does not explain why outcomes differ. The code reports confusion counts by site and exposes the denominator for each rate. The decision threshold defines what an alert means.
Read rates with their denominators
False-negative rate divides missed handoffs without alerts by all actual misses. False-positive rate divides false alarms by all actual non-misses. A site with three positive outcomes cannot support the same confidence as a site with three thousand. Report raw counts, missing labels and the share eligible for evaluation. Prevalence can differ across sites, changing precision even at similar error rates.
Audit measurement and exposure
A group gap may reflect scanner reliability, outcome definition, dispatch capacity or model behavior. Do not assume a single cause from the table. Intervention itself can change subsequent labels, making historical outcomes an imperfect counterfactual for a new policy. Review feature availability, sampling and target construction before changing thresholds. Feature timing and measurement error are separate checks.
Avoid declaring fairness from one metric
Matching alert rates, false-negative rates and false-positive rates are different policy objectives; they need not agree when underlying event rates differ. Decide which harms and constraints matter to affected groups, then assess the tradeoffs with stakeholders and the relevant governance team. The table is diagnostic evidence, not a legal determination or a certificate of fairness. A model can pass an aggregate metric while failing a small intersectional slice.
Recheck after deployment
Store group definitions, model version, threshold and outcome maturity rule with each report. Watch both error and review burden after a policy change, because a fixed threshold can drift when site mix changes. The monitoring lesson supplies the delayed-label clock. The project combines subgroup error with escalation capacity.
Implementation
# Site, actual miss, alert issued. Rows are mature future outcomes.
cases = [
("Harbor", 1, 1), ("Harbor", 1, 0), ("Harbor", 0, 1),
("Harbor", 0, 0), ("Inland", 1, 1), ("Inland", 1, 1),
("Inland", 0, 0), ("Inland", 0, 0),
]
def site_error_audit(rows):
sites = {}
for site, actual, alert in rows:
counts = sites.setdefault(site, {"tp": 0, "fn": 0, "fp": 0, "tn": 0})
key = ("tp" if alert else "fn") if actual else ("fp" if alert else "tn")
counts[key] += 1
report = {}
for site, counts in sites.items():
positives = counts["tp"] + counts["fn"]
negatives = counts["fp"] + counts["tn"]
report[site] = {
**counts,
"positive_support": positives,
"negative_support": negatives,
"false_negative_rate": counts["fn"] / positives if positives else None,
"false_positive_rate": counts["fp"] / negatives if negatives else None,
}
return report
audit = site_error_audit(cases)
assert audit["Harbor"]["false_negative_rate"] == 0.5
assert audit["Inland"]["false_positive_rate"] == 0
assert audit["Harbor"]["positive_support"] == 2Performance and operating cost
One pass over N mature labeled cases costs O(N) expected time and O(G) counters for G groups. Intersections increase G and shrink support. Reliable labels, protected-data handling where applicable, and review of policy effects cost more than the count operation.
Common Mistakes
- Do not call a rate with tiny support a stable group property.
- Do not equate one parity metric with a complete fairness assessment.
- Do not compare groups whose outcome clocks or eligibility rules differ.
Read next
- Selective prediction and review capacity
- Rare-event precision, recall and changing prevalence
- Decision thresholds: choose an action from probabilities and error costs
- Uncertainty and escalation review project
Continue the workflow: Federated evaluation by site and denominator.
Continue the workflow: Partial dependence, ICE curves and support limits.
