Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Group error gaps and policy audit

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A group audit computes outcomes and error rates separately for relevant populations, then investigates differences in labels, exposure and intervention cost before changing policy.

Specify the groups and action

A handoff alert sends a case to an overloaded review desk. Compare false-negative and false-positive rates by site, shift or another justified group; if people are affected, relevant protected-group analysis may also be required under the applicable governance process. A selection-rate difference alone does not explain why outcomes differ. The code reports confusion counts by site and exposes the denominator for each rate. The decision threshold defines what an alert means.

Read rates with their denominators

False-negative rate divides missed handoffs without alerts by all actual misses. False-positive rate divides false alarms by all actual non-misses. A site with three positive outcomes cannot support the same confidence as a site with three thousand. Report raw counts, missing labels and the share eligible for evaluation. Prevalence can differ across sites, changing precision even at similar error rates.

Audit measurement and exposure

A group gap may reflect scanner reliability, outcome definition, dispatch capacity or model behavior. Do not assume a single cause from the table. Intervention itself can change subsequent labels, making historical outcomes an imperfect counterfactual for a new policy. Review feature availability, sampling and target construction before changing thresholds. Feature timing and measurement error are separate checks.

Avoid declaring fairness from one metric

Matching alert rates, false-negative rates and false-positive rates are different policy objectives; they need not agree when underlying event rates differ. Decide which harms and constraints matter to affected groups, then assess the tradeoffs with stakeholders and the relevant governance team. The table is diagnostic evidence, not a legal determination or a certificate of fairness. A model can pass an aggregate metric while failing a small intersectional slice.

Recheck after deployment

Store group definitions, model version, threshold and outcome maturity rule with each report. Watch both error and review burden after a policy change, because a fixed threshold can drift when site mix changes. The monitoring lesson supplies the delayed-label clock. The project combines subgroup error with escalation capacity.

Implementation

python
# Site, actual miss, alert issued. Rows are mature future outcomes.
cases = [
    ("Harbor", 1, 1), ("Harbor", 1, 0), ("Harbor", 0, 1),
    ("Harbor", 0, 0), ("Inland", 1, 1), ("Inland", 1, 1),
    ("Inland", 0, 0), ("Inland", 0, 0),
]

def site_error_audit(rows):
    sites = {}
    for site, actual, alert in rows:
        counts = sites.setdefault(site, {"tp": 0, "fn": 0, "fp": 0, "tn": 0})
        key = ("tp" if alert else "fn") if actual else ("fp" if alert else "tn")
        counts[key] += 1
    report = {}
    for site, counts in sites.items():
        positives = counts["tp"] + counts["fn"]
        negatives = counts["fp"] + counts["tn"]
        report[site] = {
            **counts,
            "positive_support": positives,
            "negative_support": negatives,
            "false_negative_rate": counts["fn"] / positives if positives else None,
            "false_positive_rate": counts["fp"] / negatives if negatives else None,
        }
    return report

audit = site_error_audit(cases)
assert audit["Harbor"]["false_negative_rate"] == 0.5
assert audit["Inland"]["false_positive_rate"] == 0
assert audit["Harbor"]["positive_support"] == 2

Performance and operating cost

One pass over N mature labeled cases costs O(N) expected time and O(G) counters for G groups. Intersections increase G and shrink support. Reliable labels, protected-data handling where applicable, and review of policy effects cost more than the count operation.

Common Mistakes

  • Do not call a rate with tiny support a stable group property.
  • Do not equate one parity metric with a complete fairness assessment.
  • Do not compare groups whose outcome clocks or eligibility rules differ.

Read next

Continue the workflow: Federated evaluation by site and denominator.

Continue the workflow: Partial dependence, ICE curves and support limits.

ai-data
machine-learning
Storage details