A machine-learning task needs a prediction moment, a mature label, a baseline policy and an evaluation split that resembles deployment.
Machine learning starts with a baseline and a valid prediction clock
Specify what the model may know
For a new receipt submission, the model might predict whether manual review will be needed. The prediction happens when the submission arrives. A final review note is an outcome, not an input. Freeze that clock before collecting features. Feature availability] is more important than adding a sophisticated estimator to an invalid dataset.
Build a non-model baseline
Compare against sending all receipts to review, no receipts to review, and a simple rule such as “review high-value submissions.” Record review volume and missed cases. The model is useful only if it improves a decision within staffing and error-cost limits. Accuracy alone is weak when most receipts never need review.
Split by deployment conditions
A random split may repeat the same store or account patterns on both sides. Hold out groups if serving new stores; hold out future periods if serving next month. Fit imputers and encoders inside training folds. Validation design] is part of the task definition, not a final code detail.
Plan for action and monitoring
A probability requires a threshold and an owner for ambiguous cases. Version the model, feature schema and decision policy together. Track data availability, score distribution, queue load and outcomes after labels mature. Stop rollout if a critical slice regresses, even if the average score improves.
Implementation
baseline = {
"policy": "review receipts above the approved amount band",
"prediction_time": "submission accepted",
"target": "manual review required within seven days",
"unit": "one submission",
"capacity_per_day": 47,
}
assert baseline["capacity_per_day"] > 0Performance and operating cost
Baseline calculation is often O(N), while a trained model adds feature pipelines, repeated fitting and online inference. Start with the cheapest policy that meets the decision objective.
Common Mistakes
- Do not train before defining the prediction moment.
- Do not compare a model against no baseline.
- Do not call accuracy a complete deployment metric.
Read next
- Prediction-time feature availability: reject future information before training
- Group and time validation: split by the failure you expect in production
- Decision thresholds: choose an action from probabilities and error costs
- Data science begins with a decision, population and clock
Related path: Dataset grain and join cardinality: protect the unit of analysis.
Continue with: Machine learning core concepts: features, labels, validation and decisions.
Continue the workflow: Regression baselines and honest holdout metrics.
