Compare a simple triage rule and a classifier using time- and group-aware validation, a leakage-safe pipeline, and a fixed review-capacity policy.
Project: evaluate a receipt-review triage model
Write the release contract
Predict at submission time whether a receipt will require manual review within seven days. Allowed inputs are the amount, intake channel and account age as known then; final reviewer fields are prohibited. Hold out a future month for final evaluation and use store-grouped folds for development. Keep an explicit seven-day label-maturity buffer. Define the maximum reviews analysts can accept each day before choosing a threshold.
Build and compare
Start with a rule based on the approved amount band. Then fit a ColumnTransformer and classifier inside each development fold. Compare recall and precision at the same review-volume cap, plus per-store and per-channel counts. Inspect unexpected feature importance for a potential late-arriving field. Availability] and fit boundaries] are release requirements, not optional cleanup.
Acceptance tests
A store ID must never appear in both sides of a development fold. No preprocessing step may fit on the final month. A missing amount and a new channel must have defined behavior. Tune the threshold only on development predictions, then evaluate once on the future holdout. Report confusion counts, daily alert volume, label-mature case count and a simple calibration check. If the baseline is better under the fixed capacity, keep the baseline.
Operations handoff
Package the model with its preprocessing schema and threshold. Define monitoring for missing features, score shifts, review backlog and delayed labels. Add a rollback trigger and a named owner for the false-negative review. A notebook score without this handoff is not a deployable decision system.
Implementation
release_gate = {
"group_overlap": 0,
"future_holdout_used_for_tuning": False,
"daily_alert_cap": 47,
"unknown_channel_policy": "encode_as_unseen",
}
assert release_gate["group_overlap"] == 0
assert not release_gate["future_holdout_used_for_tuning"]
assert release_gate["daily_alert_cap"] > 0Performance and operating cost
Repeated grouped fitting costs K model fits per candidate. The service also pays online feature lookup and analyst review cost, which should be budgeted alongside prediction latency.
Common Mistakes
- Do not let late reviewer fields enter training.
- Do not optimize on the future holdout.
- Do not ship a probability without a capacity-aware action rule.
Read next
- Group and time validation: split by the failure you expect in production
- Leakage-safe preprocessing: fit every learned transform inside the training fold
- Decision thresholds: choose an action from probabilities and error costs
- Probability calibration: test whether risk scores mean what they say
Continue the workflow: Project: review shipment-delay and claim-risk models.
