Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Resistant summaries: median, trimmed mean and median absolute deviation

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A resistant summary describes central behavior when a few valid extremes can dominate the ordinary mean or standard deviation.

Choose a statistic for the question

The mean estimates an arithmetic average over all valid deliveries and is sensitive to long delays. The median describes a typical observed delivery, while a trimmed mean removes equal fractions from both tails after sorting. Neither makes late deliveries irrelevant. Publish the tail rate separately when the decision concerns service failures. Triage decides validity before a summary is calculated.

Use MAD carefully

Median absolute deviation is the median of absolute distances from the median. For delays of 21, 22, 23, 24 and 139 minutes, the mean is 45.8, the median is 23 and the raw MAD is 1. MAD can be zero when many values tie, so dividing by it to assign anomaly scores needs a stated fallback. A scale factor may be used under distribution assumptions; do not silently treat raw MAD as a standard deviation.

Keep units and strata

Report minutes, not a unitless number, and calculate summaries within comparable route types before pooling. A median can stay flat even while the slowest 5% deteriorates; add tail quantiles or threshold exceedance counts. Subgroup analysis matters when channel mix changes between periods. The same route may improve even if the aggregate median moves in the other direction.

Recompute a fixture

Start with the five delays above. Mark the 139-minute value as a valid overnight case and leave it in the mean and tail count. Then add a negative duration from a bad timestamp and quarantine that record. Recompute the valid mean, median and MAD; the first extreme affects the mean, while the invalid record should never enter any of the three.

Implementation

python
from statistics import mean, median

def delivery_summaries(delays_minutes):
    if not delays_minutes or any(delay < 0 for delay in delays_minutes):
        raise ValueError("valid nonnegative delays required")
    center = median(delays_minutes)
    return {"mean": mean(delays_minutes), "median": center,
            "raw_mad": median(abs(delay - center) for delay in delays_minutes)}

assert delivery_summaries([21, 22, 23, 24, 139]) == {"mean": 45.8, "median": 23, "raw_mad": 1}

Performance and operating cost

Sorting for medians costs O(N log N) time and O(N) memory in the standard-library implementation. The mean is O(N); reporting route-specific medians adds grouping storage and repeated sorts.

Common Mistakes

  • Do not delete valid tail cases just to make a summary stable.
  • Do not divide by MAD when it is zero.
  • Do not mistake a typical-case median for a tail-service measure.

Read next

ai-data
data-science
Storage details