Build a weekly defect ledger, inspect dispersion, and decide when a quality trend can be reported from a stable baseline.
Project: monitor parcel damage without confusing mix and measurement shifts
Build one count contract
A distribution center wants to publish its weekly damaged-parcel fraction. Define inspected parcels, defect criteria, unique parcel IDs, route and supplier fields, week boundaries, and how audits are selected. Repeated damage marks on one parcel count as one defective parcel for a p chart. Reconcile eligible audits and finalized labels each week before plotting. The p-chart lesson uses the actual inspected count for each week.
Freeze and inspect a baseline
Choose a prior period that was operationally stable under the same rubric. Check route and supplier mix, inspector assignment, and whether weekly sample size changed sharply. Compute a baseline fraction, subgroup-specific limits, and a prespecified signal rule. A stable chart can still sit above the customer’s tolerated damage level, which is a separate decision threshold. Do not discard bad weeks to make the baseline appear stable.
Investigate extra variation
Calculate a rough dispersion diagnostic and compare route-level fractions. If parcels from one truck move together, either group the data at an appropriate level or use limits that reflect the real dependence structure. Re-adjudicate a sample of images after any rubric change. The diagnostic lesson explains why simply widening bands is not enough. Record the date and effect of each operational intervention so a before-after claim is not assembled from incompatible periods.
Gate the quality report
The packet includes the audit frame, weekly numerators and denominators, baseline version, chart rule, route composition, dispersion review, rubric history, and investigated signals. The gate below blocks a public improvement claim if the label rubric changed without a bridge or if the baseline was chosen after the results. Passing the gate begins review, not proof that any one operational change caused improvement. Continue monitoring after release and retire baselines when their measurement contract changes.
Implementation
def parcel_quality_gate(audit):
if audit["duplicate_parcels"]:
return "hold:counting-unit"
if not audit["baseline_frozen_before_review"]:
return "hold:baseline-selection"
if audit["rubric_changed"] and not audit["rubric_bridge_complete"]:
return "hold:measurement-bridge"
if not audit["dispersion_reviewed"]:
return "hold:extra-variation"
return "review:quality-trend"
audit = {"duplicate_parcels": 0, "baseline_frozen_before_review": True,
"rubric_changed": True, "rubric_bridge_complete": False,
"dispersion_reviewed": True}
assert parcel_quality_gate(audit) == "hold:measurement-bridge"
assert parcel_quality_gate({**audit, "rubric_bridge_complete": True}) == "review:quality-trend"
Performance and operating cost
The gate is O(1); compiling weekly counts is O(n) over audited parcels with indexed IDs. Re-adjudicating disputed images and assessing route dependence are the expensive steps. A chart built on inconsistent labels is fast to render but cannot support a stable quality claim.
Common Mistakes
- Counting damage marks instead of defective parcels.
- Removing high-defect weeks from the baseline without a declared reason.
- Using a stable control chart as proof customer specifications are met.
- Ignoring a damage-rubric change that makes periods incomparable.
Read next
- P charts: monitor defect fractions with subgroup-specific denominators
- Process charts: test extra variation before blaming a weekly signal
- Rare proportions: keep interval uncertainty visible at zero and one
- Standard error and cluster bootstrap: resample the independent unit
- Reviewer agreement: inspect confusion cells before one kappa number
