A classifier score becomes an action only after a threshold and an operating policy define the costs of false positives and false negatives.
Decision thresholds: choose an action from probabilities and error costs
Keep ranking and action separate
A model estimates the chance that a receipt needs manual review. Sending every receipt with a score above 0.5 to an analyst is a convention, not a business law. If review capacity is 47 cases per day, the relevant constraint may be queue size; if missed fraud is expensive, recall at an acceptable false-alarm rate may matter more. Define who absorbs each error and how rapidly an analyst can clear work.
Tune without spending the test
Choose candidate thresholds on a development set, compute confusion counts and cost or capacity for each, and select one rule before the final holdout evaluation. Reusing the test set to optimize the threshold leaks evaluation information. Threshold choice does not repair poor ranking or probability calibration. Calibration] matters when the score is interpreted as risk rather than only sorted.
Reconcile the workflow
At the chosen threshold, measure true positives, false positives, false negatives, precision, recall and alerts per day for each important cohort. Include delayed labels and ambiguous outcomes as an explicit state. If capacity varies by day, a fixed threshold may overload the queue; a top-K policy has different fairness and operational properties and should be evaluated as such. Metric denominators] must specify which receipts were eligible.
Guard deployment
Version the threshold with the model and log the decision policy version. Monitor the alert rate and review capacity separately from eventual outcome metrics. A sudden alert-rate change may be upstream data drift, a feature outage or model behavior; rolling back the threshold without diagnosis can hide the cause.
Implementation
from sklearn.metrics import confusion_matrix
review_threshold = 0.37 # Chosen on development data, then frozen.
needs_review = validation_probabilities >= review_threshold
tn, fp, fn, tp = confusion_matrix(
validation_labels, needs_review, labels=[False, True]).ravel()
precision = tp / (tp + fp) if tp + fp else None
recall = tp / (tp + fn) if tp + fn else None
assert int(needs_review.sum()) == int(tp + fp)Performance and operating cost
Evaluating T thresholds across N validation rows costs O(TN) with a simple loop or O(N log N) after sorting scores. The dominant live cost is analyst time and delayed handling of true cases.
Common Mistakes
- Do not assume 0.5 matches the operating cost.
- Do not tune the threshold on the final test set.
- Do not report precision without alert volume and label maturity.
Read next
- Probability calibration: test whether risk scores mean what they say
- Group and time validation: split by the failure you expect in production
- Metric denominators and cohorts: make a rate reproducible
- Prediction-time feature availability: reject future information before training
Connected implementation
Continue the workflow: Retraining decisions: require a reason and a challenger comparison.
Continue the workflow: Hypothesis tests: pair the decision rule with an effect size.
Continue the workflow: Recommendation objective and interaction log: define what success means.
Continue the workflow: Decision objectives: define the action before optimizing a score.
Continue the workflow: Rare-event precision, recall and changing prevalence.
Continue the workflow: Linear margin classifiers and hinge loss.
Continue the workflow: Class weighting versus decision thresholds.
Continue the workflow: Selective prediction and review capacity.
Continue the workflow: Contextual bandit action logging and support.
Continue the workflow: Per-label thresholds and action cost.
Continue the workflow: Structured-output confidence and selective review.
Continue the workflow: Capacity-constrained treatment policies and guardrails.
