Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Seasonality and residual dependence in interrupted series

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Monthly interrupted-series residuals may retain seasonal structure or correlation across adjacent months, so model fit and uncertainty require checks beyond a visible policy break.

Inspect residuals after fitting the policy terms

Subtract the fitted segmented path from each observed monthly rate and examine what remains. A run of positive residuals after the intervention can indicate a misspecified lag or trend. A yearly cycle remaining in residuals means the model has not captured seasonality. The code calculates descriptive lag correlations; it is not an automatic test of significance or a complete diagnostic. The segmented model supplies the fitted path.

Separate repeated seasonal pattern from intervention timing

If the policy launches just before the annual peak, a step term may absorb the peak. Compare the same calendar months across years when enough history exists, and consider prespecified month effects or seasonal terms. With less than one full pre-policy cycle, the data cannot establish an annual pattern internally. Adding eleven month indicators to a short series can consume most of the information and make the level estimate unstable.

Dependence changes uncertainty

Adjacent residuals can move together because staffing and workload persist. Ordinary regression standard errors that assume independent errors may then understate uncertainty. Depending on series length and residual structure, a model may use a specified error process or a dependence-aware covariance estimate. Choosing the correction after scanning for the smallest p-value defeats its purpose. The plain Python calculation here stops at flagging structure; it does not produce an interval.

Check the outcome scale

A breach rate near zero or one has bounded support, while monthly breach counts depend on changing eligible-ticket volume. A linear rate model can predict outside the possible range. For very low event counts, a count or binomial model with an exposure denominator may better match the data. But a more complicated family does not fix an unmeasured concurrent event. The data contract preserves both rate components.

Report the diagnostics with the estimate

Include residual plots or sequences, lag definitions, coverage by calendar month and the final uncertainty method. A large lag correlation in twelve months is a prompt to inspect the series, not proof of a particular autoregressive model. Do not treat a corrected standard error as a correction for biased counterfactual trend. A suitable comparison addresses a different weakness.

Implementation

python
from math import sqrt

def lag_correlation(residuals, lag):
    if lag <= 0 or len(residuals) - lag < 3:
        raise ValueError("lag has too few paired months")
    earlier = residuals[:-lag]
    later = residuals[lag:]
    mean_earlier = sum(earlier) / len(earlier)
    mean_later = sum(later) / len(later)
    numerator = sum((left - mean_earlier) * (right - mean_later)
                    for left, right in zip(earlier, later))
    left_energy = sum((value - mean_earlier) ** 2 for value in earlier)
    right_energy = sum((value - mean_later) ** 2 for value in later)
    if left_energy == 0 or right_energy == 0:
        raise ValueError("constant residual segment")
    return numerator / sqrt(left_energy * right_energy)

model_residuals = [((month % 12) - 5.5) / 1000
                   for month in range(36)]
annual_lag = lag_correlation(model_residuals, 12)
nearby_lag = lag_correlation(model_residuals, 1)
assert annual_lag > .99
assert -1 <= nearby_lag <= 1

Performance and operating cost

For T months, one lag correlation is O(T) time and O(T) slice space. Checking L lags directly is O(L × T). Seasonal model estimation and dependence-aware uncertainty add model-specific costs; short series length, not computation, is often the limiting factor.

Common Mistakes

  • Do not infer independent errors from a plausible-looking fitted line.
  • Do not fit a full seasonal cycle with less than a cycle of baseline evidence and call it validated.
  • Do not treat adjusted uncertainty as a fix for a confounded intervention date.

Read next

Continue the workflow: Serial correlation: daily rows are not daily independent evidence.

ai-data
data-science
Storage details