A known intervention date defines one before-and-after comparison; searching many possible dates creates a selection problem.
Process shifts: separate a scheduled intervention from a searched break
Start with the event timeline
A warehouse changes its scanning procedure on a recorded Monday. If the date was fixed before examining scan latency, compare prespecified pre and post windows, while checking staffing, parcel mix and seasonality. If an analyst chose the date because the chart looked different, the ordinary two-window uncertainty understates the search. The code calculates only a descriptive mean gap after a declared split. It does not estimate a causal effect or an adjusted interval. Serial dependence further reduces independent information in daily readings.
Watch for calendar structure
Weekly peaks can resemble a break if the before window contains mostly weekends and the after window mostly weekdays. Use like-for-like days or a model that includes calendar effects, and retain the full time axis. A process can drift gradually rather than jump; forcing one step change may produce an attractive but misleading date. Inspect missing days and logging changes before interpreting a shift as an operational event. Rolling evaluation offers a way to test whether a time model generalizes beyond one split.
Treat searched dates as exploration
If no intervention date was planned, a scan over candidate breaks can be useful to locate investigation periods. But the best-looking split is selected from many trials. A confirmatory claim needs a separate later window, a search-aware uncertainty method or a preregistered next intervention. Report every eligible candidate date and the selection rule, not only the maximum contrast. Multiplicity explains why unadjusted repeated looks make false alarms more likely.
Keep detection and attribution apart
An apparent rise may coincide with a software update, new parcel mix or a delayed data feed. A shift detector can say the stream changed under its assumptions; it cannot identify the cause from the same series alone. Pair operational logs with time-ordered checks and state which units were independently observed. The CUSUM lesson covers prospective alarms; the project joins them to a release packet.
Implementation
from statistics import mean
def scheduled_gap(daily_latency_ms, intervention_index):
if not 2 <= intervention_index <= len(daily_latency_ms) - 2:
raise ValueError("at least two days on each side required")
before = daily_latency_ms[:intervention_index]
after = daily_latency_ms[intervention_index:]
return mean(after) - mean(before)
assert scheduled_gap([81, 83, 82, 95, 97, 96], 3) == 14
Performance and operating cost
One known split takes O(n) time and O(1) extra space over an existing list. Scanning every possible split naively costs O(n squared); prefix sums reduce the arithmetic to O(n), but they do not remove the statistical selection cost. More efficient search is not stronger causal evidence.
Common Mistakes
- Treating a visually chosen break date as if it had been prespecified.
- Calling any mean gap an effect of the warehouse procedure.
- Comparing different weekday mixes without a calendar check.
- Using daily row count as the number of independent shifts when serial correlation is present.
Read next
- CUSUM alarms: set reference and threshold before reading the stream
- Project: investigate a warehouse scan-latency alarm
- Serial correlation: daily rows are not daily independent evidence
- Multiple comparisons and peeking: protect a predeclared decision rule
- Rolling-origin backtests: make every forecast from information available then
Continue the workflow: Process capability: check stability before interpreting Cpk.
