Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Node HTTP Drain, Readiness, and Graceful Shutdown

Last updated: 5 Oct 20268 min read
tutorial
IntermediateBy AITrove Editorial

A process shutdown is a protocol with the load balancer and active clients, not simply a call to process.exit. On a termination signal, a service should report that it is no longer ready, stop accepting new connections, allow in-flight requests a bounded interval to finish, and close database or queue resources. Node's HTTP server close behavior varies by version, especially around idle keep-alive connections, so the runtime version belongs in the deployment contract. Upgraded sockets and background jobs may need separate drains. A hard deadline exists because a request can hang forever, but force-closing it can interrupt a write after commit and before response. Idempotent operation replay makes that ambiguity recoverable.

Working case

A new permit API release starts while an old instance handles a 1.8-second approval. The old process exits immediately on SIGTERM. The database commits, the client sees a reset connection, and the reviewer retries with a new operation ID, creating a second notification. In another deployment, the process calls server.close but a long-lived upgraded socket keeps the instance alive beyond the platform deadline. The repair marks readiness false, waits for traffic removal, stops accepting new HTTP work, drains normal requests within a 29-second budget, and separately tracks upgraded sessions. The approval endpoint uses a stable operation ID so a lost response can be replayed safely.

Implementation boundary

javascript
import express from 'express';
import { permitRepository } from './services.js';

const app = express();
let ready = true;
let draining = false;
app.get('/ready', (request, response) => response.sendStatus(ready ? 200 : 503));
const server = app.listen(4_781);

function beginDrain() {
  if (draining) return;
  draining = true;
  ready = false;
  const deadline = setTimeout(() => server.closeAllConnections(), 29_000);
  deadline.unref();
  server.close(async error => {
    clearTimeout(deadline);
    if (error) process.exitCode = 1;
    try { await permitRepository.close(); }
    catch { process.exitCode = 1; }
  });
}
process.on('SIGTERM', beginDrain);

Separate readiness from liveness. A draining process can be alive but should no longer receive new routed traffic. Coordinate the load balancer's observation delay with the application's admission gate; readiness changing alone does not instantly stop packets already in flight. Call server.close once, guard repeated signals, and set a bounded force-close timer that is shorter than the platform's kill deadline. Close database pools only after in-flight handlers finish or after the hard deadline is reached. Record active requests and upgraded connections; WebSocket sessions need their own close handshake and maximum wait. Treat a failure while closing a resource as a nonzero exit condition. Exercise the exact Node version and proxy behavior used in production.

Cost and boundaries

A longer drain window preserves more requests but keeps old capacity and connections alive during a rollout. A shorter window releases resources quickly yet increases ambiguous client failures. In-flight request tracking uses O(r) state for r active requests; long-lived sockets can dominate memory even after normal HTTP admission stops. A process that waits forever can stall deployment. The force-close timer can cut off an already committed write's response, so write endpoints need replay semantics. Measure requests completed during drain, forced closes, retry outcomes, time until load balancer removal, and the overlap capacity needed during deployment.

Failure trace

Send SIGTERM during a slow authorized read and verify readiness changes before connection admission closes. The read may finish inside the budget. Send SIGTERM just after an approval commits but before its response; replay with the same operation ID and verify one logical write. Send SIGTERM twice and confirm shutdown begins only once. Hold an idle keep-alive connection and an upgraded socket; neither should make the process exceed its deadline without being tracked. Force repository close to fail and verify the exit status records failure. Run the drill behind the actual proxy because local server.close behavior alone cannot prove traffic has stopped.

Verification

  • Readiness changes before new work is admitted during drain.
  • Normal requests finish or are force-closed within the deadline.
  • Repository resources close and lost writes can be replayed safely.

Practice drill

Create a readiness route and a drain controller around an HTTP server. Count active requests and record when readiness turns false, when the listener closes, and when the database pool closes. Use a 29-second maximum with an earlier alert threshold. Drive one 1.8-second read, one hanging read, and one approval whose response is lost after commit. Terminate during each case and compare client-visible outcomes. Add an upgraded connection and document its separate drain. Validate the behavior under the deployed Node runtime and proxy instead of assuming every server.close version handles idle sockets identically.

Decision note

Drain under a measured deadline, then close resources; make lost mutation responses recoverable through stable operation identity.

Common Mistakes

  • Calling process.exit immediately after SIGTERM.
  • Assuming HTTP close drains upgraded sockets or all Node versions identically.
  • Closing the database before in-flight handlers complete.

Related lessons

Express Request and Process Boundaries; Express Middleware Order and Request Identity; Express 5 Async Errors and Response Contracts; Node Request Budgets, Abort, and Event Loop Fairness; Production Signals and Incident Decisions; Realtime Connection and Event Delivery.

Apply and check

Build Project: Express permit API lifecycle and review Web Development: Express request and process quiz.

web-tech
web-development
Storage details