A visual aggregate is trustworthy only when each source row, grouping key and denominator have a declared meaning.
Aggregation grain and denominator: make every plotted mark auditable
Fix the source grain
A receipt table can contain one row per receipt while a status log contains several rows per receipt. Joining them and counting rows creates a review volume that rises when the workflow emits more events. Reduce the log to one terminal outcome per receipt before joining, or build an explicit event-level chart. Join cardinality is a precondition for visual accuracy.
Choose the denominator
A chart titled “review completion rate” needs a rule for receipts that were submitted but not yet eligible, rejected at intake or still open. Keep numerator and denominator alongside the plotted rate in the data contract. When filtering by channel, recompute both counts for the filtered cohort; never filter precomputed overall rates and average them.
Control time bins
A daily chart needs one time zone, one interval convention and an explicit policy for late events. A receipt submitted at 23:58 and completed after midnight might belong to the submission-day cohort even though its completion event occurs next day. Calendar policy keeps those two clocks apart. Document whether incomplete days are omitted or shown as provisional.
Reconcile the visual
For two channels with 34 of 47 and 18 of 23 timely receipts, the combined rate is 52 of 70, about 74.3 percent. The simple average of the two channel percentages is different. Test that the dashboard total uses summed counts. Include an empty group and one duplicated receipt ID to catch silent inflation.
Implementation
def weighted_completion_rate(channel_counts):
timely = 0
eligible = 0
for timely_count, eligible_count in channel_counts:
if not 0 <= timely_count <= eligible_count:
raise ValueError("invalid channel counts")
timely += timely_count
eligible += eligible_count
return None if eligible == 0 else timely / eligiblePerformance and operating cost
Aggregating N rows by C keys is O(N) expected time and O(C) state with a hash map. Sorting labels adds O(C log C). A database group-by can move the work closer to the source, but its join and filter contracts still need inspection.
Common Mistakes
- Do not average percentages when group sizes differ.
- Do not count event rows as independent receipts.
- Do not combine submission and completion timestamps without a cohort rule.
