Connection draining stops new work from reaching a terminating instance while allowing bounded in-flight work to finish. Removing a Pod from ready endpoints does not instantly close established TCP or HTTP/2 connections, and an upstream proxy may reuse a connection until it learns the endpoint is leaving. The application, service mesh, and load balancer each have a separate timeout and shutdown behavior.
Connection draining: let in-flight work finish while new traffic moves
Operational decision
A claims API holds a long-running upload for forty seconds while a rollout replaces its Pod. Inspect EndpointSlice conditions and Pod termination state with the read-only commands during a disposable rollout. The terminating endpoint is not ready for regular traffic, but an existing connection can still carry work. On termination, stop accepting new requests, advertise shutdown to keep-alive clients where the protocol permits it, and allow the upload to finish before the process exits. Set a grace period longer than measured high-percentile in-flight duration plus proxy propagation margin, while keeping an explicit upper bound so a stuck upload cannot block rollout forever. Exercise a client that reuses one connection and a client that opens a new connection mid-rollout. Verify no completed upload is silently duplicated and that a new request reaches a healthy Pod. A readiness probe returning failure alone does not demonstrate the established-connection path.
kubectl get endpointslices -n claims -l kubernetes.io/service-name=claims-api -o yaml
kubectl get pods -n claims -l app=claims-api -o wide
kubectl describe deployment -n claims claims-apiCost and verification
A longer grace period preserves in-flight work but uses old-Pod capacity and slows replacement. A shorter period speeds rollout while increasing aborted uploads or duplicate retries. Keep connection idle timeout, request deadline, proxy drain period, and Pod termination grace consistent. Measure active requests and client errors by revision during rollout. Retries must retain the original operation key because a connection failure can happen after the server committed an effect.
Common Mistakes
- Do not assume endpoint removal closes already established connections.
- Do not set the Pod grace period shorter than the tested drain path.
- Do not retry an ambiguous write with a new operation key.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Graceful Pod shutdown: stop accepting work before exit
- Kubernetes probes: startup, readiness, and liveness
- Gateway API routing: accepted route versus working request
- Retries and timeouts: bound the cost of a failed request
Practice and check
Advanced follow-up
- HTTP/2 drain: let existing streams finish while new calls move away
- Client connection pools: retire old endpoints after a DNS or rollout change
