Skip to content
AITroveRead. Build. Understand.
Make this comfortable

SLOs and error budgets: turn reliability into a decision

Last updated: 5 Oct 20266 min read
tutorial
IntermediateBy AITrove Editorial

A service-level indicator is a measured ratio or latency distribution for a user-relevant operation. A service-level objective sets a target for that indicator over a defined window. The error budget is the allowed miss implied by the target, not a pool of incidents the team is obliged to spend. Alerting should focus on sustained budget burn and actionable user impact.

Operational decision

For a document-download endpoint, suppose 240,000 eligible requests occur in a 28-day window and the objective is 99.92 percent successful responses. The allowed failures are 192, since 240,000 times 0.0008 equals 192. If 151 eligible requests have failed, 41 failures remain under this simplified request-count model. Exclude only events specified in the SLI definition; silently dropping inconvenient failures destroys the measure. Publish the numerator, denominator, window, exclusions, and no-traffic handling together. Review a fast-burn alert for an acute outage and a slower alert for a persistent regression. Pause risky releases when the agreed policy says the budget is depleted, while keeping room for emergency fixes. The small configuration sketch records the contract, not an alerting-system API.

yaml
service: document-download
windowDays: 28
eligibleRequests: 240000
targetSuccessPercent: 99.92
allowedFailures: 192
observedFailures: 151
remainingFailures: 41

Cost and verification

A request-count SLO can hide a small group of users who fail repeatedly, so segment and inspect the distribution without inflating metric cardinality. Longer windows smooth noise but react slowly; shorter windows catch regressions yet can swing with low traffic. Keep the arithmetic and event filters reproducible. Alerting on every failed request creates noise, while waiting until all 192 failures are spent may react too late. A budget policy should state who can decide exceptions and when it is reviewed.

Common Mistakes

  • Do not change the eligible-request definition after a bad week.
  • Do not equate missing traffic with perfect availability.
  • Do not page on a threshold without an action an operator can take.

Connected lessons

Operational follow-up

Advanced follow-up

Advanced follow-up

devops
operations
Storage details