A posterior predictive check asks whether replicated data from a fitted model exhibit the feature that matters in the observed data.
Posterior predictive checks for a hidden group gap
Specify the discrepancy
A 53-case support audit contains 11 escalations: eight among 17 West cases and three among 36 East cases. A pooled model assigns one escalation probability to every case. Its overall count may fit while the observed West-minus-East rate gap, about 38.7 percentage points, remains hard to reproduce. Choose that gap as the discrepancy before inspecting simulated results. Subgroup analysis explains why a pooled average can hide a service problem.
Generate replicated audits
With a Beta(4, 36) prior, the pooled posterior after 11 events in 53 cases is Beta(15, 78). Draw a possible rate from that posterior, then simulate 17 West and 36 East binary outcomes at the same drawn rate. Repeating this creates replicated audits under the fitted pooled model. Compare each replicated group gap with the observed gap. A small tail fraction flags a mismatch; it is not a classically calibrated p-value.
Look at what the model can and cannot express
The pooled model can only create a group gap through binomial noise. If the observed gap is rarely reproduced, investigate whether queues differ in case mix, staffing, measurement or true underlying rates. A group-specific model could help, but its fit also needs checking. Shared-prior estimates are a baseline; a fitted hierarchical model would carry uncertainty in its population distribution.
Separate model criticism from forecast testing
The same 53 outcomes informed the posterior and are used in this check. That is acceptable for diagnosing a structural miss, but it does not measure future predictive accuracy. A held-out mature period is needed for that. Held-out scoring asks a different question. Report both, and do not claim that passing a predictive check proves the model is true.
Preserve simulation details
Save the audit snapshot, chosen discrepancy, posterior parameters, random seed, number of replications and observed statistic. Repeat with enough replications to make Monte Carlo noise small relative to the decision. If a few runs disagree materially, increase draws. Investigate source labels before changing a model merely to fit one unusual group count.
Implementation
from random import Random
def replicated_gap_tail(west_events, west_cases, east_events, east_cases,
prior_alpha=4, prior_beta=36, replications=3000, seed=53):
if not (0 <= west_events <= west_cases and 0 <= east_events <= east_cases):
raise ValueError("invalid group counts")
if min(west_cases, east_cases, replications) <= 0:
raise ValueError("positive group sizes and replications required")
observed_gap = west_events / west_cases - east_events / east_cases
alpha = prior_alpha + west_events + east_events
beta = prior_beta + west_cases + east_cases - west_events - east_events
generator = Random(seed)
exceedances = 0
for _ in range(replications):
pooled_rate = generator.betavariate(alpha, beta)
west_rep = sum(generator.random() < pooled_rate for _ in range(west_cases))
east_rep = sum(generator.random() < pooled_rate for _ in range(east_cases))
exceedances += west_rep / west_cases - east_rep / east_cases >= observed_gap
return exceedances / replications
tail_fraction = replicated_gap_tail(8, 17, 3, 36)
assert 0 <= tail_fraction <= 1Performance and operating cost
With S replications and N total cases, direct Bernoulli simulation costs O(SN) time and O(1) working space. Large audits can draw binomial group counts directly. The tail fraction is a diagnostic of the chosen discrepancy under this model, not a general goodness score.
Common Mistakes
- Do not call a posterior predictive tail fraction a conventional p-value.
- Do not check only the overall event count when group disparity drives the decision.
- Do not treat an in-sample check as held-out predictive validation.
