Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Prediction interval coverage by operating slice

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Interval auditing reports the fraction of later outcomes inside each issued range and the range width by site or operating condition, with support counts.

Audit issued intervals, not reconstructed ones

A clearance forecast is useful only if its range was available when dispatch acted. Store lower and upper bounds with model version and intake timestamp. When the shipment later clears, compare the outcome with those original bounds. Recomputing an interval after its outcome is known destroys the evaluation. The code groups fixed predictions by site and reports coverage, mean width and support. The conformal lesson explains one way those bounds might be created.

Keep the denominator visible

Two covered shipments out of two produce 100-percent observed coverage but very little evidence. Report the number of mature outcomes for every slice and compare with the number scored. Slow shipments may not have cleared yet, so evaluating only resolved cases can make the range appear too good. Delayed-label monitoring provides the maturity rule. The code assumes its input rows are already eligible.

Inspect width alongside hits

A site can achieve high coverage by issuing unusably wide ranges. Report width by site and by backlog band; compare the width with staffing lead time or service window. A range that crosses every action threshold may leave dispatch with no better decision than a baseline. Pair interval coverage with point error and tail underestimation. The tail guide captures one-sided harm.

Understand the global-to-local gap

A globally calibrated residual margin can under-cover an overnight site if that site has more variable clearance times. The audit identifies the gap; it does not make a small local calibration set reliable. Consider better features, a valid stratified calibration design or a broader margin, and test changes on a new period. Avoid repeatedly optimizing slice-specific widths against the final holdout.

Escalate only when the range changes an action

Define a review rule for intervals crossing a dispatch deadline or exceeding a width budget. Measure the review queue, coverage among accepted predictions and outcomes among escalated cases. High selective accuracy can be obtained by escalating nearly everything. Selective prediction shows how to report both error and workload.

Implementation

python
issued_ranges = [
    ("Harbor", 7, 5, 9), ("Harbor", 12, 7, 11),
    ("Harbor", 8, 6, 10), ("Inland", 4, 2, 6),
    ("Inland", 9, 6, 10), ("Inland", 11, 7, 10),
]

def coverage_report(rows):
    grouped = {}
    for site, actual, lower, upper in rows:
        if lower > upper:
            raise ValueError("reversed interval")
        grouped.setdefault(site, []).append((lower <= actual <= upper, upper - lower))
    return {
        site: {
            "support": len(results),
            "coverage": sum(hit for hit, _ in results) / len(results),
            "mean_width": sum(width for _, width in results) / len(results),
        }
        for site, results in grouped.items()
    }

report = coverage_report(issued_ranges)
assert report["Harbor"]["coverage"] == 2 / 3
assert report["Inland"]["coverage"] == 2 / 3
assert report["Harbor"]["support"] == 3

Performance and operating cost

Auditing M issued intervals costs O(M) expected time and O(S) grouped summary state for S slices; the teaching code stores per-row results and uses O(M) memory. Data retention, maturation checks and human review are the larger operating costs.

Common Mistakes

  • Do not recompute an issued interval using the observed outcome.
  • Do not report coverage without support and width.
  • Do not evaluate only fast-clearing cases when slow labels are pending.

Read next

ai-data
machine-learning
Storage details