Build a replayable hourly detector that distinguishes low volume, missing input and one incident that triggers repeated observations.
Project: build and audit a receipt-volume anomaly queue
Prepare the fixture
Create two tenants with different ordinary hourly volume, weekday patterns and active-session exposure. Add a real low-volume episode, one planned closure, a duplicate receipt event, a missing partition and a collector outage. Give each event an ID, event time and ingestion time. The decision contract determines which records are observations and which can become incident labels.
Build two baselines
Start with a comparable-weekday median and then test an exposure-adjusted rate with a documented denominator floor. Preserve source health and baseline version for each decision. Do not convert missing hours to zero. Report the raw count, expected count, residual and score so a reviewer can replay one alert without relying on an opaque dashboard.
Replay and score
Choose the alert threshold on an earlier period under a fixed daily investigation budget; reserve a later period for the final backtest. Group repeated low hours into one incident ticket. Report incident recall, detection delay, confirmed precision on matured labels, pending labels and daily ticket count. Slice by tenant size and record the result for the planned closure separately.
Hand off a working queue
Provide event fixture, calendar, baseline code, replay command, alert packet, triage dispositions and a source-health dashboard. Assert that the duplicate event does not raise the count, the collector outage routes to source health, and a replay does not open a second ticket. Document one threshold change with its expected queue impact.
Implementation
def distinct_receipts(event_rows):
receipt_ids = set()
for event in event_rows:
if event["type"] == "submitted":
receipt_ids.add(event["receipt_id"])
return len(receipt_ids)
assert distinct_receipts([{"type": "submitted", "receipt_id": "r-47"},
{"type": "submitted", "receipt_id": "r-47"}]) == 1Performance and operating cost
A pass over N events costs O(N) expected time and O(U) memory for U distinct receipt IDs. Durable replay safety requires a bounded or persistent dedup ledger; the in-memory set is only the project oracle for a fixed fixture.
Common Mistakes
- Do not count a retried receipt twice.
- Do not score an absent partition as a confirmed zero.
- Do not tune the alert rule on the final backtest period.
