Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CUSUM alarms: set reference and threshold before reading the stream

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A cumulative-sum monitor adds small standardized departures from an in-control target and signals when evidence crosses a fixed decision limit.

Freeze a stable baseline

A depot monitors mean scan latency per shift. Estimate its target and baseline standard deviation from a documented stable period with the same measurement procedure. For a one-sided upward chart, standardize each new shift result, subtract a reference allowance and accumulate positive excess, resetting the sum to zero when evidence runs against the shift. Choose the allowance and decision limit before the monitored series arrives. The code shows the recurrence; it does not select thresholds or calculate false-alarm rates. The break lesson separates prospective detection from an after-the-fact search.

Interpret the reference allowance

A smaller allowance can retain weaker positive deviations, while a larger one filters more noise. The decision limit controls how much accumulated evidence is needed. Both affect false alarms and detection delay, so calibrate them against an acceptable in-control alarm frequency and the shift size that matters to operations. Do not copy a threshold from a different variable with different units or serial dependence. Varying denominator charts handle a distinct fraction outcome.

Account for dependence and mix

If successive shift averages share the same outages or workload, the standardized inputs are not independent. An alarm rate calibrated under independence can be wrong. Changes in parcel mix may also raise latency without a scanner defect. Rebaseline only through a documented rule after investigation; resetting whenever the chart alarms conceals a persistent shift. Preserve the full running path, missing shifts and any replacement of baseline estimates. The serial-correlation lesson explains the information loss.

Connect alarm to an action

An alarm starts diagnosis, not a declaration that a particular software release caused delay. Record the first crossing, threshold version, affected sites, raw shift measurements and competing operational events. A later quiet period does not erase a prior alarm from the audit trail. The project gates a claim on baseline integrity, calendar checks and a response owner.

Implementation

python
def upward_cusum(shift_latency_ms, target_ms, baseline_sd, allowance, limit):
    if baseline_sd <= 0 or allowance < 0 or limit <= 0:
        raise ValueError("valid baseline and chart settings required")
    running = 0.0
    first_alarm = None
    path = []
    for shift_index, latency in enumerate(shift_latency_ms):
        standardized = (latency - target_ms) / baseline_sd
        running = max(0.0, running + standardized - allowance)
        path.append(running)
        if running >= limit and first_alarm is None:
            first_alarm = shift_index
    return first_alarm, path

alarm, chart = upward_cusum([80, 81, 86, 87, 88], 80, 4, 0.25, 2.0)
assert alarm == 3 and chart[-1] >= 2

Performance and operating cost

The monitor takes O(n) time and O(n) space when retaining n chart values; alerting with only the latest state is O(1) space. Calibration by historical simulation costs more and depends on a credible stable baseline. One fast alarm rule cannot make correlated shifts independent.

Common Mistakes

  • Tuning the decision limit after looking at the alerts.
  • Resetting the baseline after an alarm without recording the investigation.
  • Calling the first alarm the cause or exact onset of a change.
  • Applying independent-input false-alarm claims to serially correlated shift data.

Read next

ai-data
applied-statistics
Storage details