Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: check a Bayesian support-rate model before acting

Last updated: 5 Oct 20265 min read
project
IntermediateBy AITrove Editorial

Build a reproducible model audit with posterior replication, prior alternatives, future scoring and a mature-batch ledger.

Freeze the evidence

Use 53 mature audited support cases with 11 seven-day escalations. West contributes eight of 17 cases; East contributes three of 36. Keep case IDs, queue labels, close dates, maturity dates and source snapshots. A second, later holdout contains 29 mature cases with five escalations. The earlier period alone fits the model; the later period is reserved for scoring. Cohort rules must reconcile all counts.

Check prior and model structure

Record Beta(4, 36) as the working prior and Beta(1, 9) and Beta(12, 108) as weak and strong alternatives. Update each using the same 11 of 53 evidence, then compare posterior means and action probabilities at a 17% rate threshold. Simulate 17 West and 36 East outcomes from the pooled posterior and compare group-rate gaps. Posterior replication exposes a mismatch hidden by the total.

Score the future period

Use the working posterior Beta(15, 78) to predict each case in the next 29-case holdout with the frozen next-case rate 15/93. Compare binary log loss to a 10% baseline. Keep outcomes hidden until predictions are stored, then report both overall and queue-level scores. Held-out scoring measures future prediction rather than in-sample fit.

Replay the monitoring ledger

Split the 53 fit cases into a 19-case batch with three escalations and a 34-case batch with eight. Check that each batch is mature and unique, then replay the prior to Beta(15, 78). Add fixtures for an immature batch and a repeated ID; both must fail. The monitoring contract makes the state reproducible after corrections.

Write the decision record

Publish data cutoff, prior rationales, group discrepancy, sensitivity grid, holdout loss, threshold rule, staffing cost assumptions and unresolved case-mix changes. A model can forecast better than a baseline while still failing an important queue check. State which finding drives the action and what additional evidence would change it. The staffing objective remains an operational choice, not a property of Bayes’ rule.

Implementation

python
def audit_rate_model(west, east, prior=(4, 36)):
    west_events, west_cases = west
    east_events, east_cases = east
    if min(west_cases, east_cases, *prior) <= 0:
        raise ValueError("positive group sizes and prior shapes required")
    if not (0 <= west_events <= west_cases and 0 <= east_events <= east_cases):
        raise ValueError("invalid group counts")
    events = west_events + east_events
    cases = west_cases + east_cases
    return {"events": events, "cases": cases,
            "posterior": (prior[0] + events, prior[1] + cases - events),
            "observed_gap": west_events / west_cases - east_events / east_cases}

packet = audit_rate_model((8, 17), (3, 36))
assert packet["events"] == 11 and packet["cases"] == 53
assert packet["posterior"] == (15, 78)

Performance and operating cost

Group reconciliation is O(G) for G queues; the two-queue fixture is O(1). Posterior replication adds O(SN) work for S simulations and N cases, while scoring a future period is O(N). Human review of outcome maturity, prior relevance and action cost is essential.

Common Mistakes

  • Do not use holdout labels to tune the prior before scoring.
  • Do not treat an overall rate fit as proof of queue-level fit.
  • Do not publish an action without the cohort, prior and checkpoint history.

Read next

ai-data
data-science
Storage details