A threshold can change customer outcomes without changing model bytes, so it needs separate evidence, ownership and rollback.
Score thresholds are release policy, not model metadata
Keep score and action separate
A receipt model may emit a risk score while a policy routes the receipt to review, clears it or asks for more information. The model artifact and the threshold policy have different owners and release clocks. A new threshold can double review volume even when the model digest is untouched. Record both identities on each decision. The public response contract describes the route, while the operating card states where that route is permitted.
Select the operating point honestly
Tune a threshold on a development or policy-selection cohort with mature outcomes, then evaluate the frozen choice on a separate holdout. Define the cost of false review, missed urgent case and reviewer capacity before looking for a favorable cut. A score near 0.8 need not mean an 80 percent event probability unless calibration was checked. Revisit calibration by relevant population and time window before treating scores as probabilities. Mature outcome joins matter because reviewed receipts often receive labels earlier than cleared ones.
Version the whole decision rule
Store policy revision, model digest, eligibility, threshold, tie behavior, fallback route and effective time together. A comparison using different model digests or eligibility rules cannot isolate the threshold effect. Before promotion, replay the proposed rule over a frozen decision set and estimate review volume and missed-case rate for important slices. Slice gates must still hold; a strong aggregate result does not excuse a weak cohort. A policy approval should name a rollback revision, not just a new number.
Monitor the change as a production release
Start with a limited exposure, count actions by route and compare the observed rate with the pre-release estimate. Watch capacity, overrides and delayed outcomes; a threshold can look safe for hours while mature labels take weeks. If the manual queue saturates, revert the policy pointer without claiming the model was defective. The applied project runs two policy revisions against the same digest and demonstrates why policy change evidence is distinct from model promotion.
Implementation
def route_receipt(score, policy):
if not 0 <= score <= 1:
raise ValueError("score outside accepted range")
if not 0 <= policy["review_at"] <= 1:
raise ValueError("invalid review threshold")
return {"route": "review" if score >= policy["review_at"] else "clear",
"policy_revision": policy["revision"]}
policy = {"revision": "receipt-policy-r47", "review_at": 0.72}
assert route_receipt(0.72, policy)["route"] == "review"
assert route_receipt(0.71, policy)["route"] == "clear"
assert route_receipt(0.72, {**policy, "review_at": 0.81})[
"route"] == "clear"
Performance and operating cost
Routing one score takes O(1) time and space. Replaying n stored decisions over k candidate thresholds costs O(nk) without a sorted-score index; sorting once can reduce repeated count queries. The code shows a boundary rule, not a threshold-selection procedure. Review capacity and delayed labels dominate the operational cost of a policy change.
Common Mistakes
- Changing the threshold without a new policy revision.
- Tuning on the same holdout used for the final release claim.
- Interpreting an uncalibrated score as a probability.
- Rolling back model bytes when only the policy changed.
Read next
- Human overrides: keep decisions, reasons and labels distinct
- Project: release a receipt threshold with review evidence
- Model cards as operating contracts: scope, evidence and limits
- Slice quality gates when labels are sparse or delayed
- Prediction-outcome joins: evaluate only mature, matched decisions
Continue the workflow: Inference budget gates: cost, latency and quality together.
Continue the workflow: Calibrator releases: requalify thresholds after score remapping.
