A load balancer needs a rule for sending new requests to an application instance. Liveness answers whether the process should keep running; readiness answers whether it should receive traffic now. A process can be alive while warming critical state, migrating a local cache, or preparing to shut down. A shallow readiness probe may say healthy while an essential dependency is unusable, but a probe that flaps on a minor downstream delay can remove too many instances at once. Draining is a separate state: stop new admission, allow bounded in-flight work to finish, and close or redirect long-lived connections according to the protocol. A rollout is safe only when the health thresholds and drain deadline fit real request durations.
Ingress Health, Readiness, and Connection Drain
Working case
Six permit API replicas are replaced one at a time. A new replica reports process liveness before loading its case policy table, so the ingress sends approvals to it and gets errors. An old replica is killed immediately after it becomes unready, cutting off an export and 900 open socket clients. The revised release waits for a route-level readiness condition, marks the old replica draining, stops new ordinary traffic, gives short requests a completion window, and tells long-lived clients to reconnect through an active node. A global database pause is handled with bounded load shedding rather than making all replicas flap out of rotation simultaneously.
Implementation boundary
function mayReceiveNewPermitRequest(processAlive, policyLoaded, draining) {
return processAlive && policyLoaded && !draining;
}
console.log(mayReceiveNewPermitRequest(true, true, true));
// Output: falseDefine liveness with a cheap local signal and readiness with the minimum conditions needed to handle the routes assigned to that instance. Keep probe work bounded; a health endpoint must not itself exhaust database connections during an incident. Use consecutive success and failure thresholds to avoid oscillation, then measure detection time against the availability target. During deployment, switch the instance to draining before sending a termination signal. Stop new upgrades and ordinary admission, track active requests, finish those inside a defined grace period, and hand long-lived clients a documented reconnect path. Give exports and background jobs their own lifecycle if they cannot finish within that window. Set termination grace longer than the application drain deadline, and test with actual traffic rather than an idle node.
Cost and boundaries
Health probes consume small recurring work across every instance, and more sensitive thresholds detect failure faster while risking false removals. Long drain windows preserve requests but tie up old capacity and may slow deployments. Short windows free capacity yet increase interrupted exports and reconnect surges. A readiness check that depends on one shared service can mark the entire fleet unready during a partial outage, making recovery harder. Track healthy and draining instance counts, probe failure cause, request completion during drain, forced termination count, socket reconnect rate, and rollout error rate. Ensure capacity math includes one node leaving and one not yet ready.
Failure trace
Start a replica with its HTTP port open but policy data unavailable; liveness may pass while readiness must fail. Delay one optional telemetry dependency and ensure it does not remove the serving fleet. Drain a node under 63 concurrent short requests and several long exports, then check completion and explicit interruption results. Put 900 sockets on a node, drain it, and observe that reconnects spread rather than synchronizing. Fail a shared database dependency and verify health policy does not repeatedly remove every instance while the load balancer routes to nowhere. Restart a probe process and verify hysteresis avoids rapid flapping.
Verification
- Readiness rejects traffic before essential state loads.
- Draining stops new admission before termination.
- Long-lived connections have bounded reconnect behavior.
Practice drill
Write a state table for starting, ready, draining, and stopped replicas. Choose readiness conditions for the case-read and approval paths, failure and success thresholds, a 47-second app drain deadline, and a 62-second termination grace. Run a rolling replacement with one replica removed from six. Record the minimum healthy capacity, new-request admission, completed requests, forced closures, and reconnect load. Repeat with one optional dependency failed and one essential policy table missing.
Decision note
Health controls admission; draining preserves in-flight work within a measured deadline.
Common Mistakes
- Using an open TCP port as proof that business routes are ready.
- Terminating immediately after readiness turns off.
- Making one optional dependency remove the entire fleet.
Related lessons
Ingress Proxy and Upstream Contracts; Trusted Proxy Hops and Forwarded Client Identity; Host Authority Routing and Default Deny; Proxy Body Budgets, Timeouts, and Retry Ownership; Presence Expiry, Heartbeats, and Connection Drain; Production Signals and Incident Decisions.
Apply and check
Build Project: permit ingress cutover and safe retry and review Web Development: ingress and upstream contracts quiz.
