A plotted estimate is not a population fact; its uncertainty and sample unit affect how strongly a visual difference can be read.
Uncertainty on charts: show denominators and intervals beside estimates
Show the independent count
Two channels can both report a 75 percent completion rate, yet one may have 12 receipts and another 12,000. Put eligible counts near marks or in an accessible companion table. If customers submit repeated receipts, the independent count may be customers rather than rows. Cluster resampling keeps those observations together.
Choose an interval contract
Specify the estimator, interval construction, confidence level and sampling unit. Error bars from an independent-row formula are misleading when assignments or outcomes cluster by reviewer. A confidence interval describes a repeated-sampling procedure; it is not a probability distribution over the displayed fixed rate. Coverage requires a valid observation process.
Separate overlap from decision
Overlapping intervals do not by themselves establish that a difference is absent, and nonoverlap is not a universal experiment test. If the task is to compare channels, calculate and display an interval for the difference under the study design. Pair this with practical limits: a two-point change may matter only above a declared volume or cost threshold.
Handle thin groups
A filter may leave one receipt or none. Use a visible “insufficient data” state rather than a narrow-looking point with absent error bars. Check that each interval is within the metric domain and that chart tooling does not clip it without notice. Keep the raw numerator and denominator available for audit.
Implementation
def interval_mark(point_rate, lower_rate, upper_rate, eligible_count):
if eligible_count < 2:
return {"state": "insufficient_data", "eligible_count": eligible_count}
if not 0 <= lower_rate <= point_rate <= upper_rate <= 1:
raise ValueError("invalid rate interval")
return {"state": "measured", "point": point_rate,
"lower": lower_rate, "upper": upper_rate, "eligible_count": eligible_count}Performance and operating cost
Formatting M intervals is O(M) time and space. Interval estimation may be much more expensive, especially when resampling clusters; calculate it in a versioned data job and render the stored estimate with its construction metadata.
Common Mistakes
- Do not draw error bars without naming their meaning.
- Do not use receipt count as the independent sample size when customers repeat.
- Do not interpret overlapping intervals as a complete test of a difference.
