A segmented time-series model distinguishes an immediate level shift at intervention from a change in the rate of change afterward.
Segmented regression for level and slope changes
Define the two quantities
Suppose the breach rate was rising by half a percentage point per month before a September routing policy. At the first stable post-policy month, the observed series may sit below the projection of that old trend. That vertical difference is a level shift. If the post-policy path then falls faster or rises more slowly, its slope differs too. Report both in rate points per month, plus the modeled gap at an agreed horizon. The outcome clock fixes which month is first treated.
Fit on the declared windows
The teaching implementation fits a straight line separately to pre- and post-policy points and evaluates both lines at the first post-policy time index. Subtracting the extrapolated pre line from the post line there yields a level contrast; subtracting slopes yields a slope contrast. It is algebraically equivalent to a four-parameter segmented line on complete windows without extra covariates. It does not calculate uncertainty or handle autocorrelated errors.
Do not confuse a slope with a total effect
If the level falls by five rate points but the post slope is one point per month higher than the pre slope, the modeled gap narrows over time. At horizon h, the gap is level change plus h times slope change when h counts months since the first post point. State that horizon. An immediate improvement may disappear, and a small first-month change can accumulate into a large later gap.
Check extrapolation before interpretation
A linear pretrend may be a poor counterfactual if the rate approaches a floor, varies with season, or changes after an unrelated staffing event. Inspect residuals, ticket denominators and outliers rather than reporting only a coefficient. With very few post months, the post slope is unstable. Time dependence and seasonality affect uncertainty and model shape.
Keep the causal claim conditional
The calculation establishes a break relative to a linear continuation, not why the break happened. A concurrent policy could create it. A comparison series subject to the same external shocks may help, but only if it did not receive the routing rule or its spillovers. The controlled design compares breaks rather than raw levels.
Implementation
def fit_line(points):
if len(points) < 3:
raise ValueError("need at least three points per segment")
times = [month for month, _ in points]
values = [rate for _, rate in points]
mean_time = sum(times) / len(times)
mean_rate = sum(values) / len(values)
spread = sum((month - mean_time) ** 2 for month in times)
if spread == 0:
raise ValueError("no time variation")
slope = sum((month - mean_time) * (rate - mean_rate)
for month, rate in points) / spread
intercept = mean_rate - slope * mean_time
return intercept, slope
def segmented_change(monthly_rates, first_post):
pre = [(month, rate) for month, rate in monthly_rates.items()
if month < first_post]
post = [(month, rate) for month, rate in monthly_rates.items()
if month >= first_post]
pre_intercept, pre_slope = fit_line(pre)
post_intercept, post_slope = fit_line(post)
level = (post_intercept + post_slope * first_post -
(pre_intercept + pre_slope * first_post))
return level, post_slope - pre_slope
breach_rates = {month: .17 + .005 * month for month in range(1, 9)}
breach_rates.update({month: .155 + .002 * (month - 9)
for month in range(9, 15)})
level_shift, slope_shift = segmented_change(breach_rates, 9)
assert abs(level_shift - (-.06)) < 1e-12
assert abs(slope_shift - (-.003)) < 1e-12Performance and operating cost
With T monthly observations, two ordinary least-squares line fits cost O(T) time and O(T) space in this implementation because it materializes two lists. A production regression with seasonal terms and dependence-aware uncertainty has additional solver and diagnostic costs.
Common Mistakes
- Do not report a post slope alone as the policy effect.
- Do not read a fitted break as causal without checking simultaneous events.
- Do not use ordinary independent-error intervals when residuals are serially dependent.
