Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Machine learning core concepts: features, labels, validation and decisions

Last updated: 5 Oct 20265 min read
concept
BeginnerBy AITrove Editorial

A model maps available features to an estimated outcome; its value depends on the label definition, split, calibration and action policy.

Separate the four contracts

Features are inputs known at the prediction time. A label is the later outcome being estimated. A validation design measures performance on examples withheld from fitting and tuning. A decision policy turns the score into an action. Changing any one contract changes what the model means. The same estimator can be safe in one workflow and useless in another.

Understand generalization

A model can memorize store-specific habits or ingest a field created after review. Both can produce strong development scores without helping future decisions. Use group or time holdouts that reflect deployment. Keep preprocessing fitted within each fold. The pipeline contract] prevents a common class of test contamination.

Read scores as measurements

Precision asks what fraction of flagged receipts truly needed review. Recall asks what fraction of all such receipts were caught. Their denominators differ; both depend on a chosen threshold. A probability of 0.7 has an additional meaning only if calibration supports it. Calibration] and threshold selection] are distinct checks.

Plan for feedback delay

A submission may take days to receive a final label. Treat pending outcomes separately when monitoring. Watch feature completeness and queue load immediately, then evaluate outcome metrics on mature cohorts. Compare the deployed policy with its baseline, document overrides and retain a rollback path.

Implementation

python
evaluation_contract = {
    "positive_label": "requires_manual_review",
    "split": "future submissions from held-out stores",
    "primary_metric": "recall at a fixed daily review capacity",
    "label_maturity_days": 7,
}
assert evaluation_contract["label_maturity_days"] > 0

Performance and operating cost

Each extra fold and candidate model multiplies fit work. A compact, credible evaluation is preferable to a wide search whose split does not represent production.

Common Mistakes

  • Do not treat a feature present in a warehouse as available in real time.
  • Do not use one metric for every decision cost.
  • Do not count pending labels as negative outcomes.

Read next

Study next: Prediction-time feature availability: reject future information before training.

Study next: Group and time validation: split by the failure you expect in production.

Study next: Leakage-safe preprocessing: fit every learned transform inside the training fold.

Study next: Decision thresholds: choose an action from probabilities and error costs.

Study next: Probability calibration: test whether risk scores mean what they say.

Continue the workflow: Regression error slices and costly tails.

machine-learning
Storage details