A statistical tolerance bound supports a claim about a specified fraction of individual outcomes with a specified confidence level.
One-sided tolerance bounds: cover individual outcomes, not only the mean
Ask the population question
A line manager wants evidence that at least ninety percent of future pouches are below an upper weight threshold. A confidence interval for average weight cannot answer that: a tight mean can coexist with a wide upper tail. A one-sided upper tolerance bound targets the fraction of the individual-outcome distribution below a chosen bound. The required population fraction and confidence must be declared separately. Confidence intervals concern a parameter; quantiles describe where individual observations fall.
Use a simple distribution-free construction carefully
For independent draws from one stable continuous distribution, the sample maximum is an upper bound that covers at least fraction p of that distribution with confidence one minus p raised to sample size n. Solving for n shows why strong tail claims require many independent observations. The code returns the minimum n for this particular maximum-based plan. It does not compute a two-sided interval, correct for dependence or certify a changing process. With discrete tied values the continuous derivation may be conservative, but the sampling unit must still be defensible.
Distinguish the sample from a stable process
Taking thirty pouch weights consecutively from one short run may deliver far less independent information than thirty batches across ordinary conditions. A tolerance statement for future output needs a stable, representative sampling regime. Inspect shift, ingredient batch and device state; if the line drifts, one stationary population is a doubtful model. The capability lesson starts with stability. A maximum contaminated by a recording error should trigger an audit, not an automatic deletion.
Compare the bound with the fixed limit
After sampling under the declared plan, compare the observed sample maximum with the engineering upper limit; if the maximum exceeds that limit, this simple upper-bound demonstration cannot support the desired claim. If it does not, state the exact coverage and confidence attached to the design, sample size, time window and independence assumptions. More sophisticated tolerance methods can be shorter under justified distributional assumptions, but the assumption itself must be examined. The project combines the bound with process diagnostics.
Implementation
from math import ceil, log
def required_maximum_sample_size(population_fraction, confidence):
if not 0 < population_fraction < 1 or not 0 < confidence < 1:
raise ValueError("fraction and confidence must lie strictly between zero and one")
return ceil(log(1 - confidence) / log(population_fraction))
required = required_maximum_sample_size(0.90, 0.95)
assert required == 29
assert 1 - 0.90 ** required >= 0.95
assert 1 - 0.90 ** (required - 1) < 0.95
Performance and operating cost
Planning n from the closed-form expression is O(1). Measuring n independent items is O(n) work; storing their values and batch metadata costs O(n). The method is conservative for some aims, yet it makes the cost of a population-tail claim explicit.
Common Mistakes
- Using a narrow confidence interval for mean weight as proof most pouches comply.
- Calling consecutive measurements from one batch independent future-production draws.
- Choosing population fraction or confidence after seeing the maximum.
- Treating an upper bound for one tail as a two-sided guarantee.
Read next
- Process capability: check stability before interpreting Cpk
- Project: qualify a pouch-filling line against fixed weight limits
- Confidence intervals: interpret coverage and precision honestly
- Distribution summaries: report tails and define the outlier policy
- Repeatability: pool within-item variation without hiding drift
