Skip to content
AITroveRead. Build. Understand.
Make this comfortable

User Journey SLOs and Burn Alerts

Last updated: 7 Oct 20266 min read
tutorial
IntermediateBy AITrove Editorial

A service-level indicator is a measured fraction of good events among eligible events. Define good from the user action: a case save that commits and becomes visible, rather than an API gateway returning any 2xx response. A service-level objective names the target fraction over a window. Its error budget is the permitted bad fraction for that window. An alert should ask whether failures are spending that budget quickly enough to require action; a short and a long window together reduce noise from brief spikes while still catching sustained damage. Low-volume paths need special care because one failed request can dominate a short sample.

Working case

Reviewers save 4,700 cases in a rolling window. A new release returns 202 for every save, but 47 accepted operations never reach a final result. The API success graph appears healthy while the completed-save indicator falls. With a 99 percent target, 47 failures consume the entire allowed budget for those 4,700 events. The team pages on rapid budget spend only when both a short and a longer window show a real pattern. A separate ticket-level signal catches slow degradation. They segment by release and route but avoid creating one metric series per case ID.

Implementation boundary

javascript
function budgetState(total, bad, allowedBadBasisPoints) {
  const allowance = Math.floor(total * allowedBadBasisPoints / 10000);
  return { allowance, spent: bad, over: bad > allowance };
}
console.log(JSON.stringify(budgetState(4700, 47, 100)));
// Output: {"allowance":47,"spent":47,"over":false}

The arithmetic illustrates the budget, not an alert rule. Record each eligible user operation once, including those accepted asynchronously, and define when it becomes good, bad, or still pending. Pending work must not be counted as good merely because it has not timed out yet. Choose a deadline that reflects the task. Keep the numerator and denominator aligned across browser, API, and worker signals, and deduplicate retries under one operation ID. Maintain separate indicators for latency and correctness if a fast wrong result is possible. Build alerts from rates over measured windows, include a minimum-volume rule, and send each alert to an owner with a runbook action.

Cost and boundaries

Computing a counter ratio is O(1) per event and O(k) storage for k bounded label combinations. Adding case ID, full URL, or customer ID as a label can make k grow with traffic and raise storage and query cost sharply. A longer window smooths noise but delays detection; a shorter window catches a fast outage but can page on tiny samples. Backfilling late job outcomes can revise a recent window, so the measurement pipeline needs a documented cutoff and correction policy. Compare alert detections with actual failed user journeys during drills rather than tuning only for attractive graphs.

Failure trace

A dashboard divides HTTP 500s by all requests, including image fetches and health probes. The case-save job fails after a successful 202 response, so the dashboard reports near-perfect health. Define events at the user operation boundary, and track the accepted job to its terminal state. Another mistake pages whenever any one-minute sample has an error, waking responders for an isolated failure in a low-volume route. Require meaningful budget consumption across appropriate windows, with a distinct urgent rule for complete outages. Avoid suppressing every small route; small cohorts still need investigation paths.

Verification

  • Accepted but unfinished jobs are not counted as successful completion.
  • Retries do not inflate the eligible-operation denominator.
  • Metric labels have a bounded set of values and alerts name an owner.

Practice drill

Simulate 4,700 case saves with 47 terminal failures and check that the allowed budget is exactly consumed within numeric tolerance. Send 47 retries for one operation and verify they count as one user journey. Delay job completion past its deadline and confirm it is not called good. Add a one-minute burst of two failures in a low-volume path and compare the short and long alert windows. Then hold a moderate failure rate for hours and verify a slower alert or ticket appears before the full window is exhausted.

Decision note

Measure a finished user task and alert on sustained, meaningful loss of its permitted failure budget.

Common Mistakes

  • Equating HTTP acceptance with completed user success.
  • Adding case IDs as metric labels.
  • Paging from a tiny single-window sample without an action threshold.

Connected lessons

Production Signals and Incident Decisions; Trace Context Across Requests and Jobs; Telemetry Shapes, Redaction, and Cardinality; Incident Containment and Evidence Timeline; Accepted Operations and Status Resources; Field and Lab Performance Evidence; Load Tests and Capacity Budgets.

Apply and check

Build Project: case-save incident evidence and review Web Development: operations and sync decisions quiz.

web-tech
web-development
Storage details