Bagging trains separate predictors on replacement samples and combines their outputs; an out-of-bag prediction uses only predictors whose sample excluded that row.
Bootstrap bagging and out-of-bag evaluation
Make replacement explicit
A carrier forecasts clearance hours from the backlog seen at intake. Each bag draws as many training shipments as the training set contains, with replacement. A shipment may appear several times in one bag and not at all in another. Fit one tree or other base learner per bag, then average regression predictions. The benefit is lower sensitivity to a particular training sample when the learner is unstable. Tree leaf support remains a separate constraint.
Score a row only with trees that missed it
Out-of-bag evaluation gathers predictions for shipment i from the bags that did not contain i. A tree trained on that shipment cannot contribute to its out-of-bag score. The code uses short, fixed-threshold stumps so the sampling and score boundary are visible. Its threshold is declared, not optimized. With too few bags, some rows may have no eligible votes; that is a missing estimate, not a zero-error prediction.
Keep the estimate in its lane
Out-of-bag error is useful while developing a bagged model, but it samples the same historical period. It does not simulate a new carrier, a warehouse policy change or a future outage. If multiple scans belong to one shipment, row-level bootstrap can put related scans on both sides of the evaluation. Sample at the independent shipment or site unit when that is the deployment unit. Time and group validation checks a different question.
Compare with a simpler model
Averaging many predictors adds memory and serving work. Report the gain over a training-only median and a single constrained tree on the same future rows. OOB predictions should be stored with vote counts so sparse coverage is visible. If the ensemble only fixes one tiny slice, inspect that slice and its label quality before adopting a more expensive service. The baseline guide defines the paired comparison.
Audit reproducibility
Record the seed, sampling unit, feature order, stump settings and training cutoff. Seeds make a run repeatable, but they do not remove sampling uncertainty. Refit with several declared seeds during development, then freeze one reviewed configuration. The later forest lesson adds randomness in feature selection, which changes both diversity and how individual splits can be interpreted.
Implementation
from random import Random
shipments = [(8, 3), (11, 4), (15, 4), (19, 5),
(23, 7), (27, 8), (31, 9), (36, 11)]
def fit_backlog_stump(sampled_rows, cutoff=21):
overall = sum(hours for _, hours in sampled_rows) / len(sampled_rows)
low = [hours for backlog, hours in sampled_rows if backlog < cutoff]
high = [hours for backlog, hours in sampled_rows if backlog >= cutoff]
return (sum(low) / len(low) if low else overall,
sum(high) / len(high) if high else overall)
randomizer = Random(47)
eligible_votes = [[] for _ in shipments]
for _ in range(160):
drawn = [randomizer.randrange(len(shipments)) for _ in shipments]
low_hours, high_hours = fit_backlog_stump([shipments[index] for index in drawn])
included = set(drawn)
for index, (backlog, _) in enumerate(shipments):
if index not in included:
eligible_votes[index].append(low_hours if backlog < 21 else high_hours)
assert all(eligible_votes)
out_of_bag_mae = sum(
abs(actual - sum(votes) / len(votes))
for (_, actual), votes in zip(shipments, eligible_votes)
) / len(shipments)
assert out_of_bag_mae >= 0Performance and operating cost
For B bags of N rows, drawing and scoring these fixed stumps costs O(BN) time and O(BN) temporary vote storage; a real tree adds split-search cost and tree memory. OOB scoring reuses fitted bags but never replaces a future-period evaluation.
Common Mistakes
- Do not let a tree vote on a row it trained on when computing OOB error.
- Do not bootstrap dependent scans as if each were an independent shipment.
- Do not call an OOB score a future deployment estimate.
Read next
- Random forest feature subsampling and leaf support
- Decision-tree splits and minimum leaf support
- Group and time validation: split by the failure you expect in production
- Ensemble release review project
Continue the workflow: Random forest feature subsampling and leaf support.
Continue the workflow: Ensemble disagreement and outcome noise.
