Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Distribution summaries: report tails and define the outlier policy

Last updated: 7 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A mean alone can hide a skewed process; quantiles, missingness and an explicit outlier rule show what the data contains.

Inspect the shape first

Receipt-review delay can have a low median and a long tail from tickets held over a weekend. Report count, missing count, median and upper quantiles beside the mean. A histogram or empirical distribution can reveal modes from different workflows. Exploration] belongs on the development partition when the summary will guide model decisions.

Separate invalid from extreme

A negative delay may signal a clock or join error; a 19-day delay may be valid but operationally important. Quarantine impossible records under a documented rule. Do not trim valid extremes merely because they make a dashboard look worse. If a capped metric is useful, present the uncapped distribution and the cap count too.

Choose a stable comparison

When comparing two periods, keep the same eligibility, units and time-zone boundaries. A change in channel mix can shift the overall mean while each channel remains stable. Show per-channel counts and summaries before attributing the movement to a process change. The sampling frame] can change without any underlying service change.

Check the calculation

Use a tiny fixture with delays of 3, 4, 5, 6 and 47 hours. Its mean is 13 hours and median is 5 hours, so either value alone tells a different story. Verify the quantile method used by the implementation and keep it fixed across reports. Record missing and excluded values separately.

Implementation

python
from statistics import mean, median

def review_delay_summary(delays_hours):
    valid = [value for value in delays_hours if value is not None and value >= 0]
    rejected = sum(value is not None and value < 0 for value in delays_hours)
    if not valid:
        raise ValueError("no valid review delays")
    ordered = sorted(valid)
    upper_index = min(len(ordered) - 1, int(0.9 * len(ordered)))
    return {"count": len(valid), "missing": delays_hours.count(None),
            "rejected": rejected, "mean": mean(valid),
            "median": median(valid), "upper_observation": ordered[upper_index]}

Performance and operating cost

Sorting N delays for quantiles costs O(N log N) time and O(N) memory; streaming mean and count use O(1) memory. Approximate quantiles save memory but need an error contract.

Common Mistakes

  • Do not discard a valid tail value because it is inconvenient.
  • Do not compare means across different channel mixes without showing composition.
  • Do not hide rejected records inside a missing count.

Read next

Continue the workflow: Distribution plots: keep the tail, missing values and bin policy visible.

Continue the workflow: Regression diagnostics: residual pattern, scale and influential routes.

Continue the workflow: Independent rank comparisons: report the pair probability, not a median claim.

Continue the workflow: Conditional quantiles: optimize tail predictions with pinball loss.

ai-data
applied-statistics
Storage details