Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Beta-binomial updating with an auditable case count

Last updated: 6 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A beta prior and binary audit outcomes produce a beta posterior whose parameters add observed event and non-event counts.

Freeze the likelihood

Assume each audited eligible case has a binary seven-day escalation outcome and, conditional on one stable rate, the cases are exchangeable. If cases cluster by customer or the rate changes during the audit, a single binomial likelihood is a convenience that needs checking. An audit of 47 cases with six escalations contains 41 non-escalations. The numerator and denominator must come from the same eligible cohort. Cohort rules prevent a silent denominator change.

Apply the conjugate update

Starting with Beta(3, 27), six events and 41 non-events yield Beta(9, 68). The posterior mean is 9/77, about 11.7%; the raw audit share is 6/47, about 12.8%. The difference reflects the prior. It is not a correction for selection bias, label error or late outcome arrival. Those need separate analysis.

Explain effective prior size without a fable

The shape sum, 30, controls the weight of the prior mean in the posterior mean: 30/77 times 10%, plus 47/77 times 6/47. This algebra can be described as an effective prior sample size under this model, but the shapes are distribution parameters rather than 30 actual past observations. If prior evidence comes from an overlapping sample, adding the same cases again double-counts them.

Check sequential consistency

Process the first audit batch and then the second, or combine their event and non-event totals before updating. Both paths give the same Beta(9, 68) posterior when the prior is used once and the likelihood assumptions hold. This makes a useful data-pipeline assertion. Save batch IDs so retries do not update the posterior twice. Snapshot lineage makes a later correction traceable.

Look beyond the posterior mean

Two audits can share a similar posterior mean while having very different uncertainty because their sample sizes differ. Report the full posterior or an interval with the mean. Check whether the audit could plausibly be generated under the model. If queues have different rates, group-level analysis may be more informative than a single pooled rate.

Implementation

python
def update_beta_rate(prior_alpha, prior_beta, escalated_cases, audited_cases):
    if prior_alpha <= 0 or prior_beta <= 0:
        raise ValueError("prior shapes must be positive")
    if not 0 <= escalated_cases <= audited_cases:
        raise ValueError("events must fit within audited cases")
    return prior_alpha + escalated_cases, prior_beta + audited_cases - escalated_cases

first = update_beta_rate(3, 27, 2, 17)
second = update_beta_rate(*first, 4, 30)
assert second == update_beta_rate(3, 27, 6, 47) == (9, 68)
assert round(second[0] / sum(second), 3) == 0.117

Performance and operating cost

One update takes O(1) time and space. Updating B batches costs O(B). The important operational cost is preserving unique batch identity and outcome maturity; duplicate or premature labels change the posterior even though the arithmetic remains valid.

Common Mistakes

  • Do not add the same audit batch twice.
  • Do not treat a posterior mean as a correction for missing or mislabeled outcomes.
  • Do not claim a stable common rate when case mix or time materially changes it.

Read next

Continue the workflow: Sequential posterior monitoring with mature outcomes.

Continue the workflow: Posterior predictive checks: can one rate reproduce the group pattern?.

ai-data
data-science
Storage details