Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Decision thresholds: choose an action from probabilities and error costs

Last updated: 7 Oct 20265 min read
tutorial
BeginnerBy AITrove Editorial

A classifier score becomes an action only after a threshold and an operating policy define the costs of false positives and false negatives.

Keep ranking and action separate

A model estimates the chance that a receipt needs manual review. Sending every receipt with a score above 0.5 to an analyst is a convention, not a business law. If review capacity is 47 cases per day, the relevant constraint may be queue size; if missed fraud is expensive, recall at an acceptable false-alarm rate may matter more. Define who absorbs each error and how rapidly an analyst can clear work.

Tune without spending the test

Choose candidate thresholds on a development set, compute confusion counts and cost or capacity for each, and select one rule before the final holdout evaluation. Reusing the test set to optimize the threshold leaks evaluation information. Threshold choice does not repair poor ranking or probability calibration. Calibration] matters when the score is interpreted as risk rather than only sorted.

Reconcile the workflow

At the chosen threshold, measure true positives, false positives, false negatives, precision, recall and alerts per day for each important cohort. Include delayed labels and ambiguous outcomes as an explicit state. If capacity varies by day, a fixed threshold may overload the queue; a top-K policy has different fairness and operational properties and should be evaluated as such. Metric denominators] must specify which receipts were eligible.

Guard deployment

Version the threshold with the model and log the decision policy version. Monitor the alert rate and review capacity separately from eventual outcome metrics. A sudden alert-rate change may be upstream data drift, a feature outage or model behavior; rolling back the threshold without diagnosis can hide the cause.

Implementation

python
from sklearn.metrics import confusion_matrix

review_threshold = 0.37  # Chosen on development data, then frozen.
needs_review = validation_probabilities >= review_threshold
tn, fp, fn, tp = confusion_matrix(
    validation_labels, needs_review, labels=[False, True]).ravel()
precision = tp / (tp + fp) if tp + fp else None
recall = tp / (tp + fn) if tp + fn else None
assert int(needs_review.sum()) == int(tp + fp)

Performance and operating cost

Evaluating T thresholds across N validation rows costs O(TN) with a simple loop or O(N log N) after sorting scores. The dominant live cost is analyst time and delayed handling of true cases.

Common Mistakes

  • Do not assume 0.5 matches the operating cost.
  • Do not tune the threshold on the final test set.
  • Do not report precision without alert volume and label maturity.

Read next

Connected implementation

Continue the workflow: Retraining decisions: require a reason and a challenger comparison.

Continue the workflow: Hypothesis tests: pair the decision rule with an effect size.

Continue the workflow: Recommendation objective and interaction log: define what success means.

Continue the workflow: Decision objectives: define the action before optimizing a score.

Continue the workflow: Rare-event precision, recall and changing prevalence.

Continue the workflow: Linear margin classifiers and hinge loss.

Continue the workflow: Class weighting versus decision thresholds.

Continue the workflow: Selective prediction and review capacity.

Continue the workflow: Contextual bandit action logging and support.

Continue the workflow: Per-label thresholds and action cost.

Continue the workflow: Structured-output confidence and selective review.

Continue the workflow: Capacity-constrained treatment policies and guardrails.

machine-learning
decision-threshold-and-cost
Storage details