Each multi-label score can have its own decision threshold when missed conditions and unnecessary reviews have different costs.
Per-label thresholds and action cost
Start with calibrated score meaning
A seal-breach score and address-mismatch score need not share the same prevalence or cost. A single 0.5 cutoff may route too many address cases while missing costly seal failures. Estimate thresholds on a development cohort using a declared cost for false alarms and misses. If scores are used as probabilities, check their calibration per label first. Calibration must be assessed on cases with mature, known labels.
Select without touching the final test
The code evaluates a small grid of thresholds for each label using its known development rows. It chooses the lowest cost and reports the chosen cutoff. This is a teaching selector, not a claim that a small grid or six rows is enough for deployment. Reserve a later test cohort for the policy’s actual error and workload. The untouched-test guide treats a threshold as a tuned choice.
Check combined workload
Per-label cutoffs can each look affordable while their union sends too many parcels to a shared review desk. Deduplicate case-level actions and measure how many unique parcels are escalated. If one flag blocks dispatch while another asks for routine inspection, define precedence. Review capacity is a case-level constraint, not the sum of separate label metrics.
Keep unknown outcomes out of cost fitting
Only labels that were actually assessed belong in that label’s threshold calculation. A parcel can contribute to seal-breach selection and remain unknown for water damage. Report eligible support by label and site so a low estimated cost is not mistaken for strong evidence. Unknown-state handling makes this explicit.
Watch policy interactions
A dispatch stop caused by one predicted flag can prevent later observation of another condition. After deployment, outcome labels may depend on the action itself. Log the issued labels, final action and inspection pathway; sample a fixed audit cohort where possible. Group error auditing and the project review cover those consequences.
Implementation
development = {
"seal-breach": [(0.12, 0), (0.37, 1), (0.61, 1), (0.84, 1)],
"address-mismatch": [(0.14, 0), (0.42, 0), (0.67, 1), (0.79, 1)],
}
costs = {"seal-breach": {"miss": 7, "false_alarm": 1},
"address-mismatch": {"miss": 2, "false_alarm": 3}}
candidates = (0.25, 0.50, 0.75)
def choose_cutoffs(rows_by_label, cost_by_label, cutoffs):
selected = {}
for label, rows in rows_by_label.items():
if not rows:
raise ValueError("missing development labels")
def total_cost(cutoff):
return sum(cost_by_label[label]["miss"] if score < cutoff and truth else
cost_by_label[label]["false_alarm"] if score >= cutoff and not truth else 0
for score, truth in rows)
selected[label] = min(cutoffs, key=lambda cutoff: (total_cost(cutoff), cutoff))
return selected
cutoffs = choose_cutoffs(development, costs, candidates)
assert cutoffs == {"seal-breach": 0.25, "address-mismatch": 0.50}Performance and operating cost
Searching T candidate thresholds over N labeled rows for each of L labels costs O(LTN) in the simple implementation and O(L) output space. Sorting scores can reduce repeated scans. Operational review volume and delayed labels dominate the cost of selecting a numeric cutoff.
Common Mistakes
- Do not tune thresholds on the final test.
- Do not ignore the union of per-label alerts hitting one review queue.
- Do not fit cost on labels that were never inspected.
Read next
- Multi-label schema and unknown targets
- Multi-label metrics and denominators
- Multi-label calibration and label cardinality
- Multi-label parcel review project
Continue the workflow: Affine quantization and range audit.
