Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Posterior predictive checks for a hidden group gap

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A posterior predictive check asks whether replicated data from a fitted model exhibit the feature that matters in the observed data.

Specify the discrepancy

A 53-case support audit contains 11 escalations: eight among 17 West cases and three among 36 East cases. A pooled model assigns one escalation probability to every case. Its overall count may fit while the observed West-minus-East rate gap, about 38.7 percentage points, remains hard to reproduce. Choose that gap as the discrepancy before inspecting simulated results. Subgroup analysis explains why a pooled average can hide a service problem.

Generate replicated audits

With a Beta(4, 36) prior, the pooled posterior after 11 events in 53 cases is Beta(15, 78). Draw a possible rate from that posterior, then simulate 17 West and 36 East binary outcomes at the same drawn rate. Repeating this creates replicated audits under the fitted pooled model. Compare each replicated group gap with the observed gap. A small tail fraction flags a mismatch; it is not a classically calibrated p-value.

Look at what the model can and cannot express

The pooled model can only create a group gap through binomial noise. If the observed gap is rarely reproduced, investigate whether queues differ in case mix, staffing, measurement or true underlying rates. A group-specific model could help, but its fit also needs checking. Shared-prior estimates are a baseline; a fitted hierarchical model would carry uncertainty in its population distribution.

Separate model criticism from forecast testing

The same 53 outcomes informed the posterior and are used in this check. That is acceptable for diagnosing a structural miss, but it does not measure future predictive accuracy. A held-out mature period is needed for that. Held-out scoring asks a different question. Report both, and do not claim that passing a predictive check proves the model is true.

Preserve simulation details

Save the audit snapshot, chosen discrepancy, posterior parameters, random seed, number of replications and observed statistic. Repeat with enough replications to make Monte Carlo noise small relative to the decision. If a few runs disagree materially, increase draws. Investigate source labels before changing a model merely to fit one unusual group count.

Implementation

python
from random import Random

def replicated_gap_tail(west_events, west_cases, east_events, east_cases,
                        prior_alpha=4, prior_beta=36, replications=3000, seed=53):
    if not (0 <= west_events <= west_cases and 0 <= east_events <= east_cases):
        raise ValueError("invalid group counts")
    if min(west_cases, east_cases, replications) <= 0:
        raise ValueError("positive group sizes and replications required")
    observed_gap = west_events / west_cases - east_events / east_cases
    alpha = prior_alpha + west_events + east_events
    beta = prior_beta + west_cases + east_cases - west_events - east_events
    generator = Random(seed)
    exceedances = 0
    for _ in range(replications):
        pooled_rate = generator.betavariate(alpha, beta)
        west_rep = sum(generator.random() < pooled_rate for _ in range(west_cases))
        east_rep = sum(generator.random() < pooled_rate for _ in range(east_cases))
        exceedances += west_rep / west_cases - east_rep / east_cases >= observed_gap
    return exceedances / replications

tail_fraction = replicated_gap_tail(8, 17, 3, 36)
assert 0 <= tail_fraction <= 1

Performance and operating cost

With S replications and N total cases, direct Bernoulli simulation costs O(SN) time and O(1) working space. Large audits can draw binomial group counts directly. The tail fraction is a diagnostic of the chosen discrepancy under this model, not a general goodness score.

Common Mistakes

  • Do not call a posterior predictive tail fraction a conventional p-value.
  • Do not check only the overall event count when group disparity drives the decision.
  • Do not treat an in-sample check as held-out predictive validation.

Read next

ai-data
data-science
Storage details