A learning curve compares training and validation performance as the permitted training sample grows, exposing whether collecting more similar data is likely to help.
Learning curves and training-size diagnosis
Fix the question before drawing the curve
A carrier wants to know whether another month of labeled clearance events would help a backlog predictor. Refit the same model design on increasing chronological prefixes, then score each fit on a later development-validation period. Keep features, target construction and loss definition fixed. The code uses one-feature least squares so the fit can be inspected without a dependency; the concept applies to larger estimators. Time-aware validation keeps the clock sensible.
Read both error traces
If training and validation errors are both high and close, the model may be too simple or the available features weak. If training error is low while validation error remains much higher, the model may be too sensitive to the sample. More comparable data can sometimes narrow that gap, but only if the data-generation process is stable. A flat validation curve is evidence about this setup, not a universal statement that data collection is useless.
Keep the validation cohort independent
Growing training windows can approach a fixed validation window, but must not include it. A later validation cohort avoids learning from the future; multiple rolling cutoffs give a more dependable picture than a single date. If each new prefix includes a new warehouse or revised measurement system, the curve mixes sample size with distribution change. Annotate those events before attributing any improvement to N alone.
Avoid selecting on the final test
A curve is part of model development. If the team chooses a training horizon after reading these validation errors, use a separate final period for the release claim. Do not repeatedly redraw the curve with alternate feature sets against the same final holdout and call the last score untouched. The selection protocol explains the distinction.
Turn the shape into a decision
Compare additional label cost against the observed reduction in validation error and the operational cost of mistakes. If the gap is concentrated in one site, targeted labeling may be more useful than indiscriminate growth. If both curves are poor, investigate feature availability, target quality or model capacity before buying labels. The project asks for a specific data or modeling action backed by the curve.
Implementation
from math import isclose
history = [(8, 3), (12, 4), (16, 4), (20, 6),
(24, 7), (28, 8), (32, 10), (36, 11)]
later_validation = [(14, 4), (26, 8), (34, 10)]
def fit_backlog_line(rows):
mean_backlog = sum(backlog for backlog, _ in rows) / len(rows)
mean_hours = sum(hours for _, hours in rows) / len(rows)
centered = [(backlog - mean_backlog, hours - mean_hours)
for backlog, hours in rows]
denominator = sum(backlog ** 2 for backlog, _ in centered)
if denominator == 0:
raise ValueError("no training backlog variation")
slope = sum(backlog * hours for backlog, hours in centered) / denominator
return mean_hours - slope * mean_backlog, slope
def mae(rows, intercept, slope):
return sum(abs(hours - (intercept + slope * backlog))
for backlog, hours in rows) / len(rows)
curve = []
for training_size in (4, 6, 8):
training_prefix = history[:training_size]
intercept, slope = fit_backlog_line(training_prefix)
curve.append((training_size, mae(training_prefix, intercept, slope),
mae(later_validation, intercept, slope)))
assert [point[0] for point in curve] == [4, 6, 8]
assert all(point[1] >= 0 and point[2] >= 0 for point in curve)
assert isclose(curve[-1][1], mae(history, *fit_backlog_line(history)))Performance and operating cost
For K training sizes and a one-feature closed-form fit, the direct implementation costs O(KN) time and O(K) output space. Cross-validated curves multiply fits by fold count; model-specific fit cost can dominate. Use a small, declared grid of sample sizes before scheduling expensive ensembles.
Common Mistakes
- Do not let a training prefix overlap its validation cohort.
- Do not infer that more data will help when its population or labeling process differs.
- Do not use a final test curve to choose model complexity or training horizon.
Read next
- Gradient boosting residuals and early stopping
- Paired bootstrap intervals for model gain
- Group and time validation: split by the failure you expect in production
- Ensemble release review project
Continue the workflow: Active-learning evaluation and stopping by value.
