Build a time-ordered latency monitor and release a shift claim only after baseline, calendar and logging checks.
Project: investigate a warehouse scan-latency alarm
Reconcile the series
A network of warehouses reports a sudden increase in time from barcode scan to inventory acknowledgement. Define that clock precisely, pin the event schema version and count missing or retried scans separately. Aggregate at the shift or another justified unit, retaining workload, site, software version and timestamp. An application change that moves the clock start can imitate slower processing even if customers see no difference. Measurement drift is therefore part of the incident review.
Choose the analysis path
If the rollout date was recorded before the series was inspected, run the planned before-and-after diagnostic with matched weekdays and workload checks. If the date was found from the plot, label the result exploratory and validate on later shifts. For continuing monitoring, freeze a stable baseline and prospective CUSUM settings. The break lesson prevents a searched split from masquerading as a planned comparison; the chart lesson sets the alarm mechanism.
Investigate competing explanations
Trace release logs, staffing, parcel size, scanner model, backlog and data-feed delay. Compare the direction and timing of changes across sites that adopted the release on different dates, but do not claim random assignment where rollout was chosen by site readiness. Keep the entire alarm path and report how serial dependence affects the uncertainty around a change. A significant before-after contrast cannot by itself distinguish software from concurrent operations changes.
Deliver the incident packet
The packet includes the frozen metric and baseline, threshold configuration, complete time series, alarm time, planned or searched status, calendar comparison, missingness and competing-event log. The gate below stops a causal-looking claim when a logging definition changed or the alert was tuned on the same series. Once the packet passes, an incident owner can decide whether to roll back, run a targeted experiment or monitor further.
Implementation
def scan_shift_gate(review):
if not review["clock_definition_stable"]:
return "hold:metric-definition"
if not review["baseline_frozen_before_monitoring"]:
return "hold:baseline"
if not review["calendar_and_load_checked"]:
return "hold:calendar-load"
if review["searched_date_reported_as_planned"]:
return "hold:search-disclosure"
return "review:shift-incident"
review = {"clock_definition_stable": True, "baseline_frozen_before_monitoring": True,
"calendar_and_load_checked": False, "searched_date_reported_as_planned": False}
assert scan_shift_gate(review) == "hold:calendar-load"
assert scan_shift_gate({**review, "calendar_and_load_checked": True}) == "review:shift-incident"
Performance and operating cost
The gate is O(1). Aggregating n scan events by shift is O(n) expected time, and preserving the event trail uses O(n) space. Chart computation is cheap; validating the metric and comparing operational histories dominate the work. A smaller algorithm does not replace a stable clock.
Common Mistakes
- Treating a software rollout date found after looking at latency as prespecified.
- Claiming a CUSUM alarm proves a particular release caused the shift.
- Dropping missing or retried scans from the denominator without disclosure.
- Rebaselining silently to make the alarm disappear.
Read next
- Process shifts: separate a scheduled intervention from a searched break
- CUSUM alarms: set reference and threshold before reading the stream
- Serial correlation: daily rows are not daily independent evidence
- Process charts: test extra variation before blaming a weekly signal
- Blocked randomization inference: permute labels only inside assignment blocks
