A classic Prometheus histogram exports cumulative bucket counters, a count, and a sum. A bucket whose upper bound matches the latency objective can count requests that met it. Per-process summary quantiles cannot be averaged into a correct service-wide percentile. The bucket layout is part of the instrumentation contract: once a threshold is missing, a dashboard cannot reconstruct its exact count from adjacent buckets.
Latency histograms: put a bucket at the actual SLO threshold
Operational decision
A receipt API promises that successful requests finish within 470 milliseconds. Instrument one duration histogram on every serving instance with a bucket at 0.47 seconds and a bounded outcome label. The query below divides fast successful observations by all observed requests, so errors also consume the request budget. Confirm the metric actually has the stated bucket and outcome labels before using it. Compare the result with synthetic requests and access-log counts over the same five-minute range; missing instances and scrape failures can bias both numerator and denominator. During a canary, aggregate the same metric by revision to expose a slow cohort before relying on the fleet total. Keep a separate traffic-volume guard so a service with almost no observations does not look confidently healthy. If the objective changes, update instrumentation and rollout compatibility before switching the alert expression.
sum(rate(receipt_request_duration_seconds_bucket{le="0.47",outcome="success"}[5m])) / sum(rate(receipt_request_duration_seconds_count[5m]))Cost and verification
Each extra classic bucket creates more time series for every label combination, raising scrape, storage, and query cost. Too few buckets hide the objective boundary; too many compound a cardinality problem. A five-minute rate reacts quickly but becomes noisy for low traffic. The fraction is meaningful only if every request enters the count and the success label is set consistently. Compare counts across replicas and test a known slow request rather than trusting a smooth percentile graph.
Common Mistakes
- Do not average precomputed per-instance quantiles to report fleet latency.
- Do not claim an exact threshold fraction when no matching bucket was emitted.
- Do not exclude errors from both sides of a user-facing request SLO without stating that policy.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- SLOs and error budgets: turn reliability into a decision
- Metric cardinality: keep observability usable during a surge
- Canary analysis: compare a small cohort without hiding its failures
- Scrape staleness: separate a failed target from a missing target
