Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Metric denominators and cohorts: make a rate reproducible

Last updated: 5 Oct 20265 min read
tutorial
BeginnerBy AITrove Editorial

A rate is a numerator divided by a defined eligible population over a defined clock; changing either side changes the question.

Write the eligibility rule first

A team reports “receipts reviewed within 24 hours.” Does the denominator include receipts submitted late yesterday, canceled submissions, and receipts still awaiting review? Define the cohort by submission time in UTC, identify one row per submission, and state how pending records are treated. If a submission has not yet had 24 hours to mature, exclude it from a finalized cohort or mark it pending; counting it as late at the cutoff biases the latest day downward. Grain checks] prevent review-event multiplication.

Keep time boundaries explicit

Use a half-open submission interval [start, end). Convert source timestamps to a single timezone before comparing elapsed time; daylight-saving changes otherwise turn “one calendar day” into a different duration. Decide whether the metric uses event time or ingestion time. A correction arriving after the report cutoff should trigger a documented restatement, not silently change a published number.

Expose the components

Publish numerator, denominator, pending count, missing outcome count and extraction timestamp beside the rate. Compare the same cohort in a later extract to identify late-arriving records. A rate of 93% across 14 receipts has different uncertainty from 93% across thousands. An interval] can describe sampling variability, but it will not repair an incorrect denominator or delayed event feed.

Implementation

python
cutoff = pd.Timestamp("2026-09-27T12:00:00Z")
start = pd.Timestamp("2026-09-20T00:00:00Z")
end = pd.Timestamp("2026-09-27T00:00:00Z")
cohort = submissions.loc[
    submissions["submitted_at"].ge(start) &
    submissions["submitted_at"].lt(end) &
    submissions["submitted_at"].le(cutoff - pd.Timedelta(hours=24))]
on_time = cohort["reviewed_at"].sub(cohort["submitted_at"]).le(
    pd.Timedelta(hours=24)).fillna(False)
rate = on_time.sum() / len(cohort) if len(cohort) else None

Performance and operating cost

After timestamps are parsed, cohort filtering and rate computation are O(N) time and O(N) temporary mask space. The harder cost is maintaining consistent event-time and late-arrival policy across reports.

Common Mistakes

  • Do not divide by review events when the unit is submissions.
  • Do not finalize a cohort before every row has had the allowed observation window.
  • Do not hide missing outcomes inside a single percentage.

Read next

Connected implementation

Continue the workflow: Population, estimand and sampling frame: name the quantity before calculating.

Continue the workflow: Aggregation grain and denominator: make every plotted mark auditable.

Continue the workflow: Subgroup distributions and aggregation reversal.

Continue the workflow: Cohort entry and survivorship bias: count the cases that could fail.

Continue the workflow: SQL conditional aggregation: rates with an auditable denominator.

Continue the workflow: Retention cohorts: return behavior only after a full observation window.

Continue the workflow: Restricted mean unresolved time at a business horizon.

Continue the workflow: Control limits and service targets answer different questions.

data-science
metric-denominator-and-cohort
Storage details