Metrics aggregate measurements over time, logs record discrete events, and traces connect work across components. None substitutes for the others. A useful dashboard starts with a user-facing operation and its latency, error, and traffic signals, then lets an operator follow a failing request toward the responsible service or dependency.
Observability: join metrics, logs, and traces
Operational decision
A receipt API records a request duration histogram and a result counter, emits structured logs with a trace identifier, and propagates trace context to the document service. Name operations by stable route templates rather than raw customer IDs. Keep sensitive payloads out of telemetry. During a partial failure, compare the affected route and deployment version, then inspect a sampled trace for the slow span. The PromQL sketch calculates five-minute error ratio using only requests classified as server failures; adapt the metric names and status labels to the actual instrumentation. Also watch client-visible failures outside HTTP 5xx, such as invalid responses or queue delays. A dashboard that shows a green CPU graph while customers cannot receive receipts has the wrong primary signal.
sum(rate(receipt_requests_total{result="server_error"}[5m]))
/
sum(rate(receipt_requests_total[5m]))Cost and verification
High-cardinality labels such as request IDs or email addresses can make metric storage costly and slow. Put per-request identifiers in logs and traces instead, with retention and access rules. Sampling lowers trace cost but can miss rare failures; tune it with the incident needs of the service. A zero denominator produces an absent or invalid ratio, which is not proof of health. Pair alert rules with minimum traffic and an explicit no-data decision.
Common Mistakes
- Do not put personal identifiers in metric labels.
- Do not treat no data as zero errors.
- Do not instrument infrastructure while omitting user-visible operations.
Connected lessons
- DevOps Tutorial
- Kubernetes probes: startup, readiness, and liveness
- SLOs and error budgets: turn reliability into a decision
- Incident response: contain impact, then learn
Operational follow-up
Advanced follow-up
Advanced follow-up
Advanced follow-up
Advanced follow-up
Advanced follow-up
- Scrape staleness: separate a failed target from a missing target
- Remote-write backlog: budget the gap between local samples and long-term storage
Memory failure follow-up
- Cgroup memory signals: read reclaim before the OOM counter
- Memory growth: distinguish retained data from useful cache
Related Prompt Engineering lesson: Prompt telemetry: measure failures without copying private payloads.
