Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Quantile crossing and tail coverage: audit ordered forecasts by group

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Separately fitted quantile models can produce incoherent ordering, and overall coverage can conceal undercoverage in an important segment.

Check prediction ordering first

If a model reports both a median and a ninetieth-percentile dispatch time for the same booking, the higher quantile should not fall below the median. Independent fits can cross, especially where data are thin or features fall outside training support. Count crossings on the future holdout and inspect which route and service groups produce them. Sorting the predictions for display hides a model defect; if rearrangement is chosen, state that it changes the predictions and reevaluate its coverage. Pinball loss evaluates each level separately.

Measure coverage with denominators

For each quantile level, calculate the fraction of held-out actual durations at or below the predicted threshold. At a nominal 0.9 level, approximately 90 percent is a calibration target for a well-specified future population, subject to sampling error. Also calculate by route, booking period, and service tier. A sample of twelve remote deliveries cannot support the same precision as thousands of local deliveries; show counts rather than treating both percentages equally.

Distinguish calibration from sharpness

A model can achieve very high coverage by predicting implausibly long times for everyone. Report pinball loss or threshold width along with coverage, and compare against a frozen simple baseline. Use a time-separated validation window and avoid tuning on the final audit period. If traffic patterns move, rolling coverage can fail even when the historical average appears acceptable. Serial dependence also matters when treating repeated weekly observations as independent evidence.

Define what happens on failure

Crossings, severe group undercoverage, missing outcomes, or a changed route mix should trigger review before service promises are published. A group with no validation support needs a fallback or restricted scope. The code audits supplied median and high-quantile predictions and computes overall high-quantile coverage; production reporting should also partition by group and attach uncertainty. The project requires that fuller decision packet.

Implementation

python
def quantile_audit(actual_minutes, median_minutes, p90_minutes):
    if not actual_minutes or not (len(actual_minutes) == len(median_minutes) ==
                                  len(p90_minutes)):
        raise ValueError("aligned validation rows required")
    crossings = sum(high < median for median, high in
                    zip(median_minutes, p90_minutes))
    coverage = sum(actual <= high for actual, high in
                   zip(actual_minutes, p90_minutes)) / len(actual_minutes)
    return {"crossings": crossings, "p90_coverage": coverage,
            "rows": len(actual_minutes)}

report = quantile_audit([24, 41, 58, 35], [27, 39, 33, 30], [40, 52, 48, 28])
assert report == {"crossings": 1, "p90_coverage": 0.5, "rows": 4}

Performance and operating cost

The audit is O(n) time and O(1) extra space for n validation rows. Group reporting adds O(g) counters for g segments. Correcting and refitting a crossed model costs more; merely sorting its two outputs is cheap but requires a fresh calibration check.

Common Mistakes

  • Silently sorting crossed quantiles and retaining the old validation claim.
  • Treating ninety-percent nominal coverage as exact in a tiny subgroup.
  • Using very wide thresholds to meet coverage without reporting loss.
  • Evaluating on records selected by whether their dispatch scan arrived.

Read next

ai-data
applied-statistics
Storage details