Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Latency histograms: put a bucket at the actual SLO threshold

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A classic Prometheus histogram exports cumulative bucket counters, a count, and a sum. A bucket whose upper bound matches the latency objective can count requests that met it. Per-process summary quantiles cannot be averaged into a correct service-wide percentile. The bucket layout is part of the instrumentation contract: once a threshold is missing, a dashboard cannot reconstruct its exact count from adjacent buckets.

Operational decision

A receipt API promises that successful requests finish within 470 milliseconds. Instrument one duration histogram on every serving instance with a bucket at 0.47 seconds and a bounded outcome label. The query below divides fast successful observations by all observed requests, so errors also consume the request budget. Confirm the metric actually has the stated bucket and outcome labels before using it. Compare the result with synthetic requests and access-log counts over the same five-minute range; missing instances and scrape failures can bias both numerator and denominator. During a canary, aggregate the same metric by revision to expose a slow cohort before relying on the fleet total. Keep a separate traffic-volume guard so a service with almost no observations does not look confidently healthy. If the objective changes, update instrumentation and rollout compatibility before switching the alert expression.

promql
sum(rate(receipt_request_duration_seconds_bucket{le="0.47",outcome="success"}[5m])) / sum(rate(receipt_request_duration_seconds_count[5m]))

Cost and verification

Each extra classic bucket creates more time series for every label combination, raising scrape, storage, and query cost. Too few buckets hide the objective boundary; too many compound a cardinality problem. A five-minute rate reacts quickly but becomes noisy for low traffic. The fraction is meaningful only if every request enters the count and the success label is set consistently. Compare counts across replicas and test a known slow request rather than trusting a smooth percentile graph.

Common Mistakes

  • Do not average precomputed per-instance quantiles to report fleet latency.
  • Do not claim an exact threshold fraction when no matching bucket was emitted.
  • Do not exclude errors from both sides of a user-facing request SLO without stating that policy.

Connected lessons

Practice and check

devops
operations
Storage details