Histograms and quantile summaries answer shape questions that a single mean or trend line cannot settle.
Distribution plots: keep the tail, missing values and bin policy visible
Distinguish populations
Review delay among completed receipts excludes receipts still open. A histogram of completed delays therefore cannot describe the full submission cohort without a censoring policy. State the population and cutoff. Distribution summaries help reveal a long tail, while a separate pending count shows records that do not yet have a delay.
Choose bins deliberately
A histogram with bins of one hour may appear noisy; a bin of one day can conceal a 24-hour service threshold. Fix bin edges across compared periods and label whether the left or right edge is included. Do not auto-bin each channel independently and then compare the resulting bar heights as though the units were identical.
Retain the tail
A rare 47-hour delay can matter more operationally than a high-volume cluster near five hours. If the axis is cropped to show the center, disclose the number above the cap and provide a full-range or quantile companion view. Invalid negative delays should be separated from valid extremes. Axis policy must remain stable across cohorts.
Make a small fixture
Start with delays of 3, 4, 5, 6 and 47 hours. The mean is 13 and the median is 5. A chart that omits the last value while still reporting the original mean is internally inconsistent. Test zero valid observations, one extreme observation and a delayed arrival that changes a closed reporting period.
Implementation
def delay_bins(delays_hours, edges):
if len(edges) < 2 or edges[0] < 0 or any(left >= right for left, right in zip(edges, edges[1:])):
raise ValueError("bin edges must increase from a nonnegative start")
counts = [0] * (len(edges) - 1)
overflow = 0
for delay in delays_hours:
if delay < edges[0]:
raise ValueError("delay falls below the declared range")
location = next((index for index, (left, right) in enumerate(zip(edges, edges[1:]))
if left <= delay < right), None)
if location is None:
overflow += 1
else:
counts[location] += 1
return counts, overflowPerformance and operating cost
This direct bin search costs O(NB) for N observations and B bins, with O(B) output space. A binary search on sorted edges reduces lookup to O(N log B). Store missing and overflow counts separately so a faster renderer cannot erase them.
Common Mistakes
- Do not compare histograms with different bin edges as if they were identical.
- Do not trim a valid operational tail without reporting the omitted count.
- Do not treat an unfinished receipt as a completed zero-hour delay.
Read next
- Axis scales, baselines and labels: prevent a correct number from telling a false story
- Uncertainty on charts: show denominators and intervals beside estimates
- Distribution summaries: report tails and define the outlier policy
- Missing data policy: distinguish absence from a measured zero
Continue the workflow: Contextual anomaly baselines: seasonality, exposure and missing periods.
Continue the workflow: Spatial aggregation: exposure, unstable cells and location privacy.
Continue the workflow: Synthetic-data utility: downstream tests on an untouched holdout.
Continue the workflow: Resistant summaries: median, trimmed mean and median absolute deviation.
