A load-balancer health check decides whether one backend receives new traffic. A TCP connect proves less than an HTTP response; an HTTP response from a sidecar can still succeed when the application is broken. Readiness may deliberately fail during startup or drain. A synthetic transaction travels through more of the real route, but it is too expensive or effectful to run at every per-instance health-check interval.
Load-balancer health: a green socket is not a working transaction
Operational decision
A checkout gateway returns a healthy TCP handshake while its receipt handler cannot reach the ledger. The load balancer continues routing because its check touches only the listener. Define a cheap per-instance readiness endpoint that confirms the process can accept a request and has its required local configuration; reserve end-to-end settlement checks for a lower-frequency synthetic monitor with disposable data. Inspect which host and path the load balancer probes, its interval and failure threshold, and whether it uses the same TLS and routing chain as clients. During a test, break the application handler while leaving the listener alive, then confirm the readiness decision changes within the planned withdrawal time and a synthetic request reports the user failure. Restore the dependency and verify successful checks are required before reentry. Avoid coupling every instance's readiness to a shared database outage if that would remove all capacity and prevent controlled recovery.
kubectl -n checkout get endpointslice -l kubernetes.io/service-name=receipt-api -o wide
kubectl -n checkout get pods -l app=receipt-api -o wideCost and verification
Deeper checks catch more faults but can load shared dependencies and create a correlated outage. A high failure threshold delays withdrawal; a low one can flap under brief network loss. Compare observed withdrawal and reentry times with request error rates, route membership, and the independent synthetic result. The two commands show Kubernetes service endpoints and Pods, but they do not reveal the external load balancer's health decision; collect that separately from the configured controller.
Common Mistakes
- Do not treat a successful TCP connect as proof that the business route works.
- Do not turn an effectful payment request into a frequent backend health probe.
- Do not assume a Pod's Ready condition exactly matches an external load balancer's target state.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes probes: startup, readiness, and liveness
- Synthetic transactions: measure the route a user actually takes
- Gateway API routing: accepted route versus working request
- Connection draining: let in-flight work finish while new traffic moves
