Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Process shifts: separate a scheduled intervention from a searched break

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A known intervention date defines one before-and-after comparison; searching many possible dates creates a selection problem.

Start with the event timeline

A warehouse changes its scanning procedure on a recorded Monday. If the date was fixed before examining scan latency, compare prespecified pre and post windows, while checking staffing, parcel mix and seasonality. If an analyst chose the date because the chart looked different, the ordinary two-window uncertainty understates the search. The code calculates only a descriptive mean gap after a declared split. It does not estimate a causal effect or an adjusted interval. Serial dependence further reduces independent information in daily readings.

Watch for calendar structure

Weekly peaks can resemble a break if the before window contains mostly weekends and the after window mostly weekdays. Use like-for-like days or a model that includes calendar effects, and retain the full time axis. A process can drift gradually rather than jump; forcing one step change may produce an attractive but misleading date. Inspect missing days and logging changes before interpreting a shift as an operational event. Rolling evaluation offers a way to test whether a time model generalizes beyond one split.

Treat searched dates as exploration

If no intervention date was planned, a scan over candidate breaks can be useful to locate investigation periods. But the best-looking split is selected from many trials. A confirmatory claim needs a separate later window, a search-aware uncertainty method or a preregistered next intervention. Report every eligible candidate date and the selection rule, not only the maximum contrast. Multiplicity explains why unadjusted repeated looks make false alarms more likely.

Keep detection and attribution apart

An apparent rise may coincide with a software update, new parcel mix or a delayed data feed. A shift detector can say the stream changed under its assumptions; it cannot identify the cause from the same series alone. Pair operational logs with time-ordered checks and state which units were independently observed. The CUSUM lesson covers prospective alarms; the project joins them to a release packet.

Implementation

python
from statistics import mean

def scheduled_gap(daily_latency_ms, intervention_index):
    if not 2 <= intervention_index <= len(daily_latency_ms) - 2:
        raise ValueError("at least two days on each side required")
    before = daily_latency_ms[:intervention_index]
    after = daily_latency_ms[intervention_index:]
    return mean(after) - mean(before)

assert scheduled_gap([81, 83, 82, 95, 97, 96], 3) == 14

Performance and operating cost

One known split takes O(n) time and O(1) extra space over an existing list. Scanning every possible split naively costs O(n squared); prefix sums reduce the arithmetic to O(n), but they do not remove the statistical selection cost. More efficient search is not stronger causal evidence.

Common Mistakes

  • Treating a visually chosen break date as if it had been prespecified.
  • Calling any mean gap an effect of the warehouse procedure.
  • Comparing different weekday mixes without a calendar check.
  • Using daily row count as the number of independent shifts when serial correlation is present.

Read next

Continue the workflow: Process capability: check stability before interpreting Cpk.

ai-data
applied-statistics
Storage details