Build a reproducible receipt-review report with explicit grain, join checks, missingness, cohort rules and a reviewable output manifest.
Project: audit a receipt-review rate from raw submissions
Input and decision
Use two small files: submissions contain submission_id, store_id, submitted_at and receipt_amount; review events contain submission_id, event_id, reviewed_at and status. Produce the weekly proportion of eligible submissions reviewed within 24 hours. The project is complete only if another analyst can reproduce the result from the same snapshot and cutoff. Do not use personal identifiers or production credentials in the sample files.
Work in contracts
Assert submission_id is unique in the submissions file. For events, define how corrections and repeated reviews are ordered before selecting the relevant review. Perform a checked join that preserves one row per submission. Define a half-open week window and exclude submissions that have not had 24 hours to mature at the cutoff. Keep missing outcomes visible. The grain lesson] and cohort lesson] provide the two central checks.
Deliverables and acceptance
Save a script, input manifest, aggregated output and a short decision note. The output must include eligible submissions, on-time reviews, late reviews, pending or unknown rows, rate and extraction cutoff. A fixture with one submission having three review events must still count once. A fixture with a missing review timestamp must not silently disappear. A changed cutoff should produce a new report version rather than overwrite the first result.
Stretch test
Split the report by store while preserving the global denominator. Compare one week with the next and explain whether the change could come from late-arriving events or a shifted store mix. Do not claim a staffing change caused an improvement without a credible comparison design.
Implementation
def verify_report(report, submissions):
assert report["eligible"] <= len(submissions)
assert report["on_time"] + report["late"] + report["unknown"] == report["eligible"]
assert 0 <= report["on_time_rate"] <= 1
assert report["snapshot_id"] and report["cutoff_utc"]
return reportPerformance and operating cost
The project is O(N + E log E) for N submissions and E review events when latest-event selection sorts history. A larger source should push reductions into storage and keep the same assertions.
Common Mistakes
- Do not count review events as submissions.
- Do not hide pending outcomes in a dropped-row count.
- Do not publish a rate without its cutoff and input snapshot.
Read next
- Dataset grain and join cardinality: protect the unit of analysis
- Missing data policy: distinguish absence from a measured zero
- Metric denominators and cohorts: make a rate reproducible
- Reproducible analysis snapshots: pin data, code and cutoff together
Related path: Project: evaluate a receipt-review triage model.
