Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Connection draining: let in-flight work finish while new traffic moves

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Connection draining stops new work from reaching a terminating instance while allowing bounded in-flight work to finish. Removing a Pod from ready endpoints does not instantly close established TCP or HTTP/2 connections, and an upstream proxy may reuse a connection until it learns the endpoint is leaving. The application, service mesh, and load balancer each have a separate timeout and shutdown behavior.

Operational decision

A claims API holds a long-running upload for forty seconds while a rollout replaces its Pod. Inspect EndpointSlice conditions and Pod termination state with the read-only commands during a disposable rollout. The terminating endpoint is not ready for regular traffic, but an existing connection can still carry work. On termination, stop accepting new requests, advertise shutdown to keep-alive clients where the protocol permits it, and allow the upload to finish before the process exits. Set a grace period longer than measured high-percentile in-flight duration plus proxy propagation margin, while keeping an explicit upper bound so a stuck upload cannot block rollout forever. Exercise a client that reuses one connection and a client that opens a new connection mid-rollout. Verify no completed upload is silently duplicated and that a new request reaches a healthy Pod. A readiness probe returning failure alone does not demonstrate the established-connection path.

bash
kubectl get endpointslices -n claims -l kubernetes.io/service-name=claims-api -o yaml
kubectl get pods -n claims -l app=claims-api -o wide
kubectl describe deployment -n claims claims-api

Cost and verification

A longer grace period preserves in-flight work but uses old-Pod capacity and slows replacement. A shorter period speeds rollout while increasing aborted uploads or duplicate retries. Keep connection idle timeout, request deadline, proxy drain period, and Pod termination grace consistent. Measure active requests and client errors by revision during rollout. Retries must retain the original operation key because a connection failure can happen after the server committed an effect.

Common Mistakes

  • Do not assume endpoint removal closes already established connections.
  • Do not set the Pod grace period shorter than the tested drain path.
  • Do not retry an ambiguous write with a new operation key.

Connected lessons

Practice and check

Advanced follow-up

systemd operating follow-up

Linux host change follow-up

devops
operations
Storage details