Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Simulation error and replication budgets

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Monte Carlo standard error measures variation from a finite number of simulated replications while holding the input model fixed.

Identify two uncertainty layers

The support-shift simulator returns a breach flag on each modeled shift. If R independent replications yield estimated breach share p, its approximate Monte Carlo standard error is the square root of p times (1 minus p) divided by R. This quantifies numerical noise from finite simulation. It says nothing about uncertainty in the urgent probability, arrival distribution or capacity estimate. The estimand lesson separates model inputs from outputs.

Budget precision before running

For an approximate 95% half-width no larger than 0.01 under the most variable possible Bernoulli result, use p times (1 minus p) at its maximum 0.25. The planning calculation rounds up 0.25 times (1.96 divided by 0.01) squared, which is 9,604 replications. This is a fixed-run planning bound for the normal approximation, not a guarantee of exact interval coverage at rare probabilities. A rare breach may need many more runs to observe enough events for a useful estimate.

Report the denominator and interval

An output such as 0.143 without a run count hides precision. Report the event count, R, estimate and a stated interval method. The simple p ± 1.96 standard-error calculation below is a diagnostic only; it can be poor near zero or one and may cross probability bounds. For publication, use a binomial interval appropriate to the event count or increase replications. A zero observed breach count does not prove a zero breach probability.

Keep the stopping rule honest

Repeatedly running batches until the estimate passes a desired threshold or the interval looks favorable changes the analysis. Fix the replication budget before inspecting the result, or use a sequential method designed for optional stopping. Increasing R after finding a large standard error can be an engineering refinement if the policy and reporting rule were specified in advance; document that decision and rerun all candidate policies under the same release protocol.

Spend compute where it changes the decision

If the probability threshold for buying flex coverage is 0.12 and the estimate is 0.47, a tiny Monte Carlo standard error is unlikely to alter the decision. If the estimate is 0.121, more simulation may matter. Compare numerical error with input-model sensitivity: rerun plausible urgent-share or demand assumptions. Scenario stress addresses model uncertainty that a large R cannot erase.

Implementation

python
from math import ceil, sqrt
from random import Random

def planned_binary_replications(target_half_width, z_score=1.96):
    if not 0 < target_half_width < 1 or z_score <= 0:
        raise ValueError("invalid precision target")
    return ceil(0.25 * (z_score / target_half_width) ** 2)

def simulated_breach_precision(replications, seed):
    if replications < 2:
        raise ValueError("at least two replications required")
    draw_stream = Random(seed)
    breaches = sum(draw_stream.random() < 0.31 for _ in range(replications))
    estimate = breaches / replications
    standard_error = sqrt(estimate * (1 - estimate) / replications)
    return breaches, estimate, standard_error

budget = planned_binary_replications(0.01)
assert budget == 9604
events, probability, mc_error = simulated_breach_precision(24000, 287)
assert 0 < events < 24000
assert 1.96 * mc_error < 0.01

Performance and operating cost

Planning the run count is O(1). Generating R independent breach flags is O(R) time and O(1) additional space when outcomes are streamed. Halving Monte Carlo standard error requires about four times as many independent replications, which can be expensive for detailed shift models.

Common Mistakes

  • Do not call Monte Carlo error the total uncertainty of a decision.
  • Do not interpret zero simulated breaches as proof of zero risk.
  • Do not keep changing the stopping rule after reading interim outcomes.

Read next

ai-data
data-science
Storage details