Audit a policy launch, reporting delay and calendar dependence before claiming a sustained reduction in late refunds.
Project: review a refund-delay change with correlated daily outcomes
Freeze the daily measure
A returns team changed its triage policy on a recorded launch date. The proposed outcome is the fraction of eligible refunds completed late each day, based on the date each refund entered the queue. Capture daily eligible counts, completion maturity, weekdays, holidays, branch coverage and reporting-feed version. A calendar day with only a few mature cases is not comparable to a fully observed day with hundreds. The frame and exposure logic prevent denominator drift.
Find two false improvements
The first chart shows fewer late refunds immediately after launch, but the new feed reports completed refunds two days later and the post period contains more weekends. Rebuild the outcome from matured cohorts, align calendar dates and model planned weekday effects. Inspect residual runs: backlog produces positive serial correlation even after calendar adjustment. A row-independent standard error is too small for the series. The dependence lesson gives an information diagnostic, not a final test.
Compare defensible uncertainty methods
If a stable residual segment remains after the predeclared mean model, evaluate a dependence-aware interval and compare it with a range of moving-block lengths. If the level change or seasonality remains in the resampled object, the bootstrap assumption fails and a model-based time-series approach or longer collection is needed. Show pre-period fit, transition lag, post-period duration and a placebo date check. A control series can help separate a system-wide refund shift from the policy. The controlled-series lesson develops that contrast.
Issue only the supported claim
Deliver the extraction version, daily counts, maturity lag, calendar model, residual diagnostics, effect estimate and intervals by method. The gate holds a result while reporting lag or post-launch coverage is unresolved. Even a stable time-aware interval does not prove the policy caused the change when other launches coincide. State the association and remaining competing explanations. Block-length sensitivity and transition analysis remain in the review packet.
Implementation
def refund_series_gate(audit, limits):
if audit["unmatured_daily_cohorts"]:
return "hold:outcome-maturity"
if audit["feed_change_unreconciled"]:
return "hold:reporting-contract"
if audit["post_days"] < limits["minimum_post_days"]:
return "hold:post-period-support"
if not audit["calendar_and_dependence_checked"]:
return "hold:serial-uncertainty"
return "review:scoped-time-series-result"
limits = {"minimum_post_days": 47}
audit = {"unmatured_daily_cohorts": 0,
"feed_change_unreconciled": True, "post_days": 63,
"calendar_and_dependence_checked": True}
assert refund_series_gate(audit, limits) == "hold:reporting-contract"
assert refund_series_gate({**audit, "feed_change_unreconciled": False},
limits) == "review:scoped-time-series-result"
Performance and operating cost
The gate is O(1); daily cohort reconstruction is at least O(n) over refund events, and B block-bootstrap draws require O(Bn) work on the prepared series. The audit is needed because a reporting-feed delay can impersonate a policy effect.
Common Mistakes
- Dividing late completions by a denominator from a different cohort.
- Treating weekends and weekdays as exchangeable raw observations.
- Claiming policy causality from one uncontrolled before/after step.
- Running a block bootstrap before resolving the reporting-feed break.
Read next
- Serial correlation: daily rows are not daily independent evidence
- Moving-block bootstrap: preserve local dependence in resampled series
- Population, estimand and sampling frame: name the quantity before calculating
- Controlled interrupted time series with a comparison series
- Transition windows and lag sensitivity for policy series
Continue the workflow: Project: validate a case-volume forecast before it sets staffing.
