An alert rule evaluates an operational signal and sends a notification when a condition persists. A page should indicate user impact or a near-term risk that requires prompt human action. A dashboard is for investigation; sending every dashboard threshold to an on-call phone creates noise and delays response to serious incidents.
Alert design: page on impact and include a first action
Operational decision
For a document-download service, alert when a meaningful volume of requests has a sustained server-error ratio above its baseline. Route the page to the team that can pause rollout or repair the service, and include a runbook with the first checks: current release digest, affected route, dependency health, and rollback compatibility. The Prometheus rule below uses a ratio and a 6-minute pending interval; it is only a starting point. Add a minimum traffic condition and an explicit no-data alert in the real rule set, and evaluate the effect against historical data before enabling paging. Test that the notification reaches the intended person and that a resolved condition clears it. A missing metrics pipeline should trigger its own monitoring path, because an absent error series is not a healthy series.
groups:
- name: document-download
rules:
- alert: DownloadServerErrors
expr: sum(rate(download_requests_total{result="server_error"}[6m])) / sum(rate(download_requests_total[6m])) > 0.018
for: 6m
labels: {severity: page, owner: documents}
annotations: {summary: 'Document download errors exceed the review threshold'}Cost and verification
A longer pending interval reduces transient pages but delays notification. A ratio without a traffic floor can be noisy at low request volume or undefined at zero; this sketch is deliberately incomplete until those cases are handled. Too many label combinations can multiply alert instances and on-call load. Keep diagnostic labels stable and bounded. Use the service's SLO and real recovery time to choose thresholds rather than copying 1.8 percent into another service.
Common Mistakes
- Do not page on a condition with no operator action.
- Do not interpret absent telemetry as zero failures.
- Do not ship a ratio alert without checking low-traffic behavior.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- SLOs and error budgets: turn reliability into a decision
- Observability: join metrics, logs, and traces
- Incident response: contain impact, then learn
