Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Observability: join metrics, logs, and traces

Last updated: 5 Oct 20266 min read
tutorial
IntermediateBy AITrove Editorial

Metrics aggregate measurements over time, logs record discrete events, and traces connect work across components. None substitutes for the others. A useful dashboard starts with a user-facing operation and its latency, error, and traffic signals, then lets an operator follow a failing request toward the responsible service or dependency.

Operational decision

A receipt API records a request duration histogram and a result counter, emits structured logs with a trace identifier, and propagates trace context to the document service. Name operations by stable route templates rather than raw customer IDs. Keep sensitive payloads out of telemetry. During a partial failure, compare the affected route and deployment version, then inspect a sampled trace for the slow span. The PromQL sketch calculates five-minute error ratio using only requests classified as server failures; adapt the metric names and status labels to the actual instrumentation. Also watch client-visible failures outside HTTP 5xx, such as invalid responses or queue delays. A dashboard that shows a green CPU graph while customers cannot receive receipts has the wrong primary signal.

promql
sum(rate(receipt_requests_total{result="server_error"}[5m]))
/
sum(rate(receipt_requests_total[5m]))

Cost and verification

High-cardinality labels such as request IDs or email addresses can make metric storage costly and slow. Put per-request identifiers in logs and traces instead, with retention and access rules. Sampling lowers trace cost but can miss rare failures; tune it with the incident needs of the service. A zero denominator produces an absent or invalid ratio, which is not proof of health. Pair alert rules with minimum traffic and an explicit no-data decision.

Common Mistakes

  • Do not put personal identifiers in metric labels.
  • Do not treat no data as zero errors.
  • Do not instrument infrastructure while omitting user-visible operations.

Connected lessons

Operational follow-up

Advanced follow-up

Advanced follow-up

Advanced follow-up

Advanced follow-up

Advanced follow-up

Memory failure follow-up

Related Prompt Engineering lesson: Prompt telemetry: measure failures without copying private payloads.

devops
operations
Storage details