Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Bootstrap bagging and out-of-bag evaluation

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Bagging trains separate predictors on replacement samples and combines their outputs; an out-of-bag prediction uses only predictors whose sample excluded that row.

Make replacement explicit

A carrier forecasts clearance hours from the backlog seen at intake. Each bag draws as many training shipments as the training set contains, with replacement. A shipment may appear several times in one bag and not at all in another. Fit one tree or other base learner per bag, then average regression predictions. The benefit is lower sensitivity to a particular training sample when the learner is unstable. Tree leaf support remains a separate constraint.

Score a row only with trees that missed it

Out-of-bag evaluation gathers predictions for shipment i from the bags that did not contain i. A tree trained on that shipment cannot contribute to its out-of-bag score. The code uses short, fixed-threshold stumps so the sampling and score boundary are visible. Its threshold is declared, not optimized. With too few bags, some rows may have no eligible votes; that is a missing estimate, not a zero-error prediction.

Keep the estimate in its lane

Out-of-bag error is useful while developing a bagged model, but it samples the same historical period. It does not simulate a new carrier, a warehouse policy change or a future outage. If multiple scans belong to one shipment, row-level bootstrap can put related scans on both sides of the evaluation. Sample at the independent shipment or site unit when that is the deployment unit. Time and group validation checks a different question.

Compare with a simpler model

Averaging many predictors adds memory and serving work. Report the gain over a training-only median and a single constrained tree on the same future rows. OOB predictions should be stored with vote counts so sparse coverage is visible. If the ensemble only fixes one tiny slice, inspect that slice and its label quality before adopting a more expensive service. The baseline guide defines the paired comparison.

Audit reproducibility

Record the seed, sampling unit, feature order, stump settings and training cutoff. Seeds make a run repeatable, but they do not remove sampling uncertainty. Refit with several declared seeds during development, then freeze one reviewed configuration. The later forest lesson adds randomness in feature selection, which changes both diversity and how individual splits can be interpreted.

Implementation

python
from random import Random

shipments = [(8, 3), (11, 4), (15, 4), (19, 5),
             (23, 7), (27, 8), (31, 9), (36, 11)]

def fit_backlog_stump(sampled_rows, cutoff=21):
    overall = sum(hours for _, hours in sampled_rows) / len(sampled_rows)
    low = [hours for backlog, hours in sampled_rows if backlog < cutoff]
    high = [hours for backlog, hours in sampled_rows if backlog >= cutoff]
    return (sum(low) / len(low) if low else overall,
            sum(high) / len(high) if high else overall)

randomizer = Random(47)
eligible_votes = [[] for _ in shipments]
for _ in range(160):
    drawn = [randomizer.randrange(len(shipments)) for _ in shipments]
    low_hours, high_hours = fit_backlog_stump([shipments[index] for index in drawn])
    included = set(drawn)
    for index, (backlog, _) in enumerate(shipments):
        if index not in included:
            eligible_votes[index].append(low_hours if backlog < 21 else high_hours)

assert all(eligible_votes)
out_of_bag_mae = sum(
    abs(actual - sum(votes) / len(votes))
    for (_, actual), votes in zip(shipments, eligible_votes)
) / len(shipments)
assert out_of_bag_mae >= 0

Performance and operating cost

For B bags of N rows, drawing and scoring these fixed stumps costs O(BN) time and O(BN) temporary vote storage; a real tree adds split-search cost and tree memory. OOB scoring reuses fitted bags but never replaces a future-period evaluation.

Common Mistakes

  • Do not let a tree vote on a row it trained on when computing OOB error.
  • Do not bootstrap dependent scans as if each were an independent shipment.
  • Do not call an OOB score a future deployment estimate.

Read next

Continue the workflow: Random forest feature subsampling and leaf support.

Continue the workflow: Ensemble disagreement and outcome noise.

ai-data
machine-learning
Storage details