Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Graceful Pod shutdown: stop accepting work before exit

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Graceful shutdown is the coordinated end of a process after a Pod is marked for termination. Kubernetes starts a grace period, runs a preStop hook when configured, and asks the container process to stop; once the deadline expires, remaining processes can be killed. The application must close listeners or mark itself unready, finish bounded work, and acknowledge only operations whose effects are durable.

Operational decision

A receipt API has requests that may take twenty seconds. Give it a measured termination window, handle the stop signal in the server, and stop taking new work before waiting for active handlers. The YAML fragment shows a Pod-level deadline, not the application signal handler. Test a rollout while a slow receipt is in progress. Record whether the client receives a response, whether the write committed, and whether a retry creates a duplicate. Some load balancers observe endpoint changes after a delay; a short fixed sleep in preStop is not a substitute for testing that delay. If a request cannot finish within the window, rely on an idempotent retry or durable queue rather than extending shutdown forever.

yaml
spec:
  terminationGracePeriodSeconds: 47
  containers:
    - name: receipt-api
      image: registry.internal/receipt-api@sha256:REPLACE_WITH_FULL_DIGEST
      ports:
        - containerPort: 8147

Cost and verification

Long grace periods slow rollouts and node maintenance, while short ones interrupt valid work. The number in the fragment is an invented starting value; measure the high percentile of in-flight work and leave room for endpoint propagation. A handler that ignores the stop signal may run until forced termination even when the manifest looks correct. Verify from outside the cluster, since a successful process exit does not prove the user transaction succeeded.

Common Mistakes

  • Do not assume a preStop sleep completes business work.
  • Do not acknowledge a queued item before its effect is durable.
  • Do not set a grace period without testing the actual application handler.

Connected lessons

Advanced follow-up

devops
resilience
Storage details