A synthetic transaction is a controlled request that exercises a service path before a customer reports failure. A network or HTTP probe can confirm reachability, TLS, response status, and latency; it cannot prove an order or payout completed unless it performs and verifies that business operation. The probe's location, credentials, data, and cleanup determine what failure it can detect.
Synthetic transactions: measure the route a user actually takes
Operational decision
A claims service exposes a health endpoint that stays green during a database write outage. Run an external HTTP probe for TLS and route availability, then a separate synthetic claim using a dedicated low-privilege test identity and a reversible or isolated record. The PromQL expression checks the external probe's success signal; the business transaction requires its own metric and verification. Place at least one probe outside the cluster and label each target by stable service and region, without customer IDs. Alert on sustained failure and attach the last successful transaction time to the incident. Expire test credentials like any other secret and clean synthetic records under a named retention rule. A probe that only calls the same internal load balancer as the application misses public DNS and edge faults.
min_over_time(probe_success{job="claims-external-http"}[5m]) == 0Cost and verification
External probes create small but recurring traffic, storage, and notification cost. A single probe location may report a local network fault as a global outage; several locations add expense and interpretation work. A test write can have real side effects if isolation is weak, so scope it to a dedicated account and verify cleanup. The query detects five minutes of continuous probe failure for each target; a complete alert also handles absent series and separates transport failure from business-operation failure.
Common Mistakes
- Do not call a green health endpoint proof of successful writes.
- Do not use production customer credentials for synthetic transactions.
- Do not alert on probe failure without checking missing metric series and probe location.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Observability: join metrics, logs, and traces
- SLO burn-rate alerts: page on budget consumption, not isolated spikes
- DNS cutovers: budget for resolver caches and mixed destinations
- Incident response: contain impact, then learn
