Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Per-label thresholds and action cost

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Each multi-label score can have its own decision threshold when missed conditions and unnecessary reviews have different costs.

Start with calibrated score meaning

A seal-breach score and address-mismatch score need not share the same prevalence or cost. A single 0.5 cutoff may route too many address cases while missing costly seal failures. Estimate thresholds on a development cohort using a declared cost for false alarms and misses. If scores are used as probabilities, check their calibration per label first. Calibration must be assessed on cases with mature, known labels.

Select without touching the final test

The code evaluates a small grid of thresholds for each label using its known development rows. It chooses the lowest cost and reports the chosen cutoff. This is a teaching selector, not a claim that a small grid or six rows is enough for deployment. Reserve a later test cohort for the policy’s actual error and workload. The untouched-test guide treats a threshold as a tuned choice.

Check combined workload

Per-label cutoffs can each look affordable while their union sends too many parcels to a shared review desk. Deduplicate case-level actions and measure how many unique parcels are escalated. If one flag blocks dispatch while another asks for routine inspection, define precedence. Review capacity is a case-level constraint, not the sum of separate label metrics.

Keep unknown outcomes out of cost fitting

Only labels that were actually assessed belong in that label’s threshold calculation. A parcel can contribute to seal-breach selection and remain unknown for water damage. Report eligible support by label and site so a low estimated cost is not mistaken for strong evidence. Unknown-state handling makes this explicit.

Watch policy interactions

A dispatch stop caused by one predicted flag can prevent later observation of another condition. After deployment, outcome labels may depend on the action itself. Log the issued labels, final action and inspection pathway; sample a fixed audit cohort where possible. Group error auditing and the project review cover those consequences.

Implementation

python
development = {
    "seal-breach": [(0.12, 0), (0.37, 1), (0.61, 1), (0.84, 1)],
    "address-mismatch": [(0.14, 0), (0.42, 0), (0.67, 1), (0.79, 1)],
}
costs = {"seal-breach": {"miss": 7, "false_alarm": 1},
         "address-mismatch": {"miss": 2, "false_alarm": 3}}
candidates = (0.25, 0.50, 0.75)

def choose_cutoffs(rows_by_label, cost_by_label, cutoffs):
    selected = {}
    for label, rows in rows_by_label.items():
        if not rows:
            raise ValueError("missing development labels")
        def total_cost(cutoff):
            return sum(cost_by_label[label]["miss"] if score < cutoff and truth else
                       cost_by_label[label]["false_alarm"] if score >= cutoff and not truth else 0
                       for score, truth in rows)
        selected[label] = min(cutoffs, key=lambda cutoff: (total_cost(cutoff), cutoff))
    return selected

cutoffs = choose_cutoffs(development, costs, candidates)
assert cutoffs == {"seal-breach": 0.25, "address-mismatch": 0.50}

Performance and operating cost

Searching T candidate thresholds over N labeled rows for each of L labels costs O(LTN) in the simple implementation and O(L) output space. Sorting scores can reduce repeated scans. Operational review volume and delayed labels dominate the cost of selecting a numeric cutoff.

Common Mistakes

  • Do not tune thresholds on the final test.
  • Do not ignore the union of per-label alerts hitting one review queue.
  • Do not fit cost on labels that were never inspected.

Read next

Continue the workflow: Affine quantization and range audit.

ai-data
machine-learning
Storage details