Multiclass evaluation reports where each true category is sent, while an action policy compares expected costs across possible interventions instead of choosing a label mechanically.
Multiclass confusion and action costs
Define mutually exclusive outcomes
A handoff can fail because of an intake scan gap, insufficient crew or a weather block. A single-label multiclass classifier assigns one primary reason under an adjudication rule; if multiple reasons may apply simultaneously, that is a multilabel task instead. Fix the label rule before training. A confusion matrix records true categories by predicted category and exposes which mistake pairs dominate. Class prevalence should accompany the matrix.
Read each class with its denominator
The code counts a small three-class future cohort and computes recall for each true class. A high overall accuracy could still miss most weather blocks if they are rare. Report support and per-class precision as well; a model can have high recall by overpredicting a category. At low counts, one case changes a rate substantially. Keep an unassigned or uncertain path where the operational process requires it rather than forcing a reason from weak evidence.
Optimize the action, not the class name
Adding crew, repairing a scanner and rerouting are different interventions with different effects and costs. If probabilities are calibrated and the cost matrix is credible, compute the expected cost of each action and choose the least costly permitted one. The lowest-cost action need not correspond to the most probable failure label. The binary cost lesson is the simpler version of this policy.
Validate probability and cost inputs
A multiclass probability vector should be nonnegative and sum to one; its class order must match the cost columns exactly. Calibration by class and by site deserves inspection before interpreting expected costs. Cost entries should reflect resource limits and who bears the consequence, not convenient invented numbers. The example uses fictional costs only to show the arithmetic. Calibration remains a prerequisite for a probability-based dispatch rule.
Test the complete decision workflow
Report confusion, per-class support and realized intervention cost on a later, untouched cohort. Compare against a simple current rule, not only another classifier. Watch for label latency and policy feedback: rerouting can change the eventual reason recorded. Monitoring must preserve the prediction-time record and the outcome clock.
Implementation
categories = ("scan_gap", "crew_shortage", "weather_block")
future_cases = [
("scan_gap", "scan_gap"), ("scan_gap", "crew_shortage"),
("crew_shortage", "crew_shortage"), ("crew_shortage", "crew_shortage"),
("crew_shortage", "scan_gap"), ("weather_block", "weather_block"),
("weather_block", "crew_shortage"),
]
confusion = {(actual, predicted): sum(
observed == actual and assigned == predicted
for observed, assigned in future_cases)
for actual in categories for predicted in categories}
recall = {category: confusion[(category, category)] /
sum(confusion[(category, prediction)] for prediction in categories)
for category in categories}
current_risk = {"scan_gap": 0.25, "crew_shortage": 0.50, "weather_block": 0.25}
action_cost = {
"repair_scanner": {"scan_gap": 1, "crew_shortage": 6, "weather_block": 4},
"add_crew": {"scan_gap": 4, "crew_shortage": 1, "weather_block": 5},
"reroute": {"scan_gap": 5, "crew_shortage": 4, "weather_block": 1},
}
expected_cost = {action: sum(current_risk[reason] * costs[reason]
for reason in categories)
for action, costs in action_cost.items()}
chosen_action = min(expected_cost, key=expected_cost.get)
assert chosen_action == "add_crew"
assert recall["weather_block"] == 0.5
assert abs(sum(current_risk.values()) - 1) < 1e-12Performance and operating cost
Counting N labeled rows is O(N + C²) time and O(C²) memory for C classes. Comparing A actions against C probabilities costs O(AC) per decision. Maintaining accurate cost inputs and representative calibration data is usually harder than the arithmetic.
Common Mistakes
- Do not force a single label when multiple causes can co-occur.
- Do not infer a rare-class recall from one or two cases without showing support.
- Do not align probabilities and cost columns by undocumented list position.
Read next
- Rare-event precision, recall and changing prevalence
- Decision thresholds: choose an action from probabilities and error costs
- Probability calibration: test whether risk scores mean what they say
- Handoff triage model review project
Continue the workflow: Multi-label schema and unknown targets.
