Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Poisson count models: use exposure offsets and test dispersion

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A Poisson log-link model turns a count predictor into a rate comparison only when exposure is represented correctly.

Model events over time at risk

A claims operations team counts duplicate submissions per branch-month. Branches operate for different numbers of monitored claim-days. A count model with a log exposure offset predicts expected events as exposure times a modeled rate; the offset coefficient is fixed at one. Do not pass both exposure and its logarithm to a library that already logs the exposure argument. The raw rate lesson establishes the denominator before regression begins.

Interpret a coefficient on the rate scale

With a log link, exponentiating a predictor coefficient gives a conditional rate ratio at fixed included covariates and comparable exposure. A rate ratio of 1.4 means the modeled rate is 40% higher for that contrast, not that every branch will have 40% more events. Expected count also scales with exposure: double at-risk claim-days and the modeled count doubles under the same rate. The small program below calculates expected counts from a prespecified model; it is not a fitting routine.

Check variance and zeros

The basic Poisson model sets conditional variance equal to conditional mean. Duplicate submissions can arrive in bursts from one system outage, producing extra variation and many zero months. Compare residual patterns and count variance within comparable exposure bands; investigate event deduplication and offsets before changing distribution families. A negative-binomial or other model may better handle excess spread, but it cannot repair wrong claim-day exposure. Residual thinking transfers to this count setting.

Keep inference and causality separate

Branch software versions, staffing and season can confound the predictor. Standard errors may need branch clustering if repeated months share shocks. Report count, exposure, modeled rate, dispersion check, calendar support and uncertainty by branch segment. The project holds a rate claim when the offset is missing or event bursts violate the planned uncertainty; a confounder graph is needed for intervention claims.

Implementation

python
from math import exp, log

def expected_duplicate_claims(log_baseline_rate, log_rate_ratio,
                              upgraded_branch, monitored_claim_days):
    if monitored_claim_days <= 0 or upgraded_branch not in (0, 1):
        raise ValueError("invalid exposure or branch indicator")
    linear_predictor = (log_baseline_rate +
                        log_rate_ratio * upgraded_branch +
                        log(monitored_claim_days))
    return exp(linear_predictor)

base = expected_duplicate_claims(log(0.07), log(1.4), 0, 200)
upgraded = expected_duplicate_claims(log(0.07), log(1.4), 1, 200)
assert round(base, 6) == 14
assert round(upgraded, 6) == 19.6

Performance and operating cost

A single prediction takes O(1) time and space; fitting a count model requires repeated scans of the training rows. Correct exposure collection, deduplication and dependence-aware uncertainty cost more than evaluating the link function. A tiny standard error from an underdispersed working assumption can be misleading when branch outages create event bursts.

Common Mistakes

  • Comparing count coefficients without an at-risk exposure offset.
  • Logging exposure twice when a fitting library already accepts raw exposure.
  • Assuming Poisson mean-equals-variance after clustered system outages.
  • Reading a conditional rate ratio as an intervention effect.

Read next

Continue the workflow: Excess zeros: separate no opportunity from a stochastic zero count.

ai-data
applied-statistics
Storage details