Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Machine learning starts with a baseline and a valid prediction clock

Last updated: 7 Oct 20265 min read
concept
BeginnerBy AITrove Editorial

A machine-learning task needs a prediction moment, a mature label, a baseline policy and an evaluation split that resembles deployment.

Specify what the model may know

For a new receipt submission, the model might predict whether manual review will be needed. The prediction happens when the submission arrives. A final review note is an outcome, not an input. Freeze that clock before collecting features. Feature availability] is more important than adding a sophisticated estimator to an invalid dataset.

Build a non-model baseline

Compare against sending all receipts to review, no receipts to review, and a simple rule such as “review high-value submissions.” Record review volume and missed cases. The model is useful only if it improves a decision within staffing and error-cost limits. Accuracy alone is weak when most receipts never need review.

Split by deployment conditions

A random split may repeat the same store or account patterns on both sides. Hold out groups if serving new stores; hold out future periods if serving next month. Fit imputers and encoders inside training folds. Validation design] is part of the task definition, not a final code detail.

Plan for action and monitoring

A probability requires a threshold and an owner for ambiguous cases. Version the model, feature schema and decision policy together. Track data availability, score distribution, queue load and outcomes after labels mature. Stop rollout if a critical slice regresses, even if the average score improves.

Implementation

python
baseline = {
    "policy": "review receipts above the approved amount band",
    "prediction_time": "submission accepted",
    "target": "manual review required within seven days",
    "unit": "one submission",
    "capacity_per_day": 47,
}
assert baseline["capacity_per_day"] > 0

Performance and operating cost

Baseline calculation is often O(N), while a trained model adds feature pipelines, repeated fitting and online inference. Start with the cheapest policy that meets the decision objective.

Common Mistakes

  • Do not train before defining the prediction moment.
  • Do not compare a model against no baseline.
  • Do not call accuracy a complete deployment metric.

Read next

Related path: Dataset grain and join cardinality: protect the unit of analysis.

Continue with: Machine learning core concepts: features, labels, validation and decisions.

Continue the workflow: Regression baselines and honest holdout metrics.

machine-learning
Storage details