A posterior predictive check simulates outcomes from a fitted model and compares a declared feature of those outcomes with observed data.
Posterior predictive checks: can one rate reproduce the group pattern?
Choose a discrepancy that matters
The helpdesk model assumes one escalation rate across two queues. Operations sees a much higher observed rate in the night queue. Before simulating, choose the absolute difference between queue rates as the discrepancy. The goal is not to certify the model with one number; it is to ask whether replicated datasets under the common-rate model commonly exhibit a gap at least as large as the real one. Prior choice determines the posterior parameter draws used in this check.
Replicate the observation process
Draw a rate from the posterior, then draw new escalation counts for each queue at its observed number of mature tickets. Compare replicated gaps with the observed gap. Preserve queue denominators and the same outcome window; otherwise the simulation tests a different design. A small replicated-tail fraction flags a pattern the model struggles to generate, but it is not a classical p-value and does not identify the cause. Shared agents, label delay or queue-specific rates are all candidates for investigation.
Check multiple failure modes separately
A gap statistic can miss an overall count mismatch or a burst concentrated on one shift. Inspect replicated total escalations, queue gaps and temporal runs, each tied to an operational concern. Do not keep inventing checks until one produces a desired pass. Use held-out or future data when the question is predictive performance outside the fitting set. Cluster dependence matters if tickets arrive in agent or customer groups; a simple independent-binomial generator does not reproduce those correlations.
Decide what follows
If a common-rate model cannot reproduce queue differences, consider a stratified or hierarchical model, or repair the event ledger if labels differ by queue. Compare model revisions using the same outcome definition and sample frame. Keep the simulated distribution and random seed in the review packet. The project holds an automatic intervention while the night queue discrepancy remains unexplained; drift checks matter after deployment.
Implementation
from random import Random
def queue_gap_check(group_counts, prior, draws=600, seed=47):
if draws < 1 or any(not 0 <= hits <= total or total == 0
for hits, total in group_counts):
raise ValueError("invalid group counts or draw count")
observed = abs(group_counts[0][0] / group_counts[0][1] -
group_counts[1][0] / group_counts[1][1])
successes = sum(hits for hits, _ in group_counts)
trials = sum(total for _, total in group_counts)
rng = Random(seed)
extreme = 0
for _ in range(draws):
rate = rng.betavariate(prior[0] + successes,
prior[1] + trials - successes)
simulated = [sum(rng.random() < rate for _ in range(total)) / total
for _, total in group_counts]
extreme += abs(simulated[0] - simulated[1]) >= observed
return extreme / draws
tail_fraction = queue_gap_check([(7, 45), (15, 39)], (2, 18))
assert 0 <= tail_fraction <= 1
Performance and operating cost
The direct simulation uses O(d × n) time for d draws and n total ticket trials, with O(1) extra space. Vectorized or count-distribution sampling can reduce runtime for large n. The check is only as credible as the observation model: a shared-rate simulation cannot validate queue-dependent labels or correlated repeat contacts.
Common Mistakes
- Calling the replicated-tail fraction a probability that the model is true.
- Changing denominators or label windows between observed and simulated counts.
- Choosing only the discrepancy that makes the model look adequate.
- Using an in-sample predictive check as a substitute for future-period validation.
Read next
- Prior sensitivity: show when limited data leave a decision exposed
- Project: review a prior-sensitive helpdesk escalation decision
- Standard error and cluster bootstrap: resample the independent unit
- Beta-binomial updating with an auditable case count
- Calibration drift: compare sensor readings with a reference
