A time series is an ordered set of observations under a fixed calendar, sampling interval and reporting cutoff.
Time-series calendar: distinguish missing periods from measured zeros
Define the clock
Receipt submissions arrive as instants while a daily operations report uses business dates. Choose a time zone and decide whether daily bins are left-closed or right-closed. Daylight-saving changes can create local days with different hour counts; do not assume every local day contains 24 hours. Aggregation grain links the calendar to the population counted in each bin.
Preserve absence
A day with no submissions is a measured zero only if ingestion was healthy and that day was in scope. An outage or absent file is missing data. Use an explicit completeness flag rather than filling every gap with zero, because a false zero changes trends, seasonality and service decisions. Missingness policy belongs in the dataset manifest.
Handle revisions
A submission may arrive late or be corrected after a daily snapshot closes. Set a watermark and label recent days provisional. Publish a new snapshot when historical counts are revised; keep the old version for a decision audit. Store event time and ingestion time separately so a late record can be located without inventing an artificial submission date.
Exercise boundary cases
Construct events at 23:59 and 00:01 around the chosen cutoff, plus a calendar day with no source file and a day with a confirmed empty file. The first two must land in different daily bins; only the confirmed empty day may become zero. Test a repeated receipt ID so a retry does not inflate volume.
Implementation
def daily_submission_counts(events, calendar_days, complete_days):
counts = {day: 0 for day in calendar_days}
seen_receipts = set()
for receipt_id, business_day in events:
if receipt_id in seen_receipts:
raise ValueError("duplicate receipt")
seen_receipts.add(receipt_id)
if business_day in counts:
counts[business_day] += 1
return {day: counts[day] if day in complete_days else None
for day in calendar_days}Performance and operating cost
A single pass over N events and D calendar days costs O(N + D) expected time and O(N + D) state with duplicate detection. If upstream uniqueness is proven, counts need only O(D) state; retain the proof and revision policy.
Common Mistakes
- Do not convert an ingestion outage into zero demand.
- Do not group UTC timestamps by a local date without an explicit time zone.
- Do not silently overwrite a closed historical snapshot.
Read next
- Lag features: prove each input existed at forecast time
- Seasonal naive forecast: establish a baseline before fitting a model
- Aggregation grain and denominator: make every plotted mark auditable
- Missing data policy: distinguish absence from a measured zero
Continue the workflow: Difference-in-differences: compare changes under a defended trend assumption.
Continue the workflow: Watermarks and late events: define when a window becomes final.
