Skip to content
AITroveRead. Build. Understand.
Make this comfortable

HTTP/2 drain: let existing streams finish while new calls move away

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

An HTTP/2 connection can carry many concurrent streams. Removing a server from new routing does not instantly remove streams already accepted on that connection. During graceful shutdown the server must stop accepting new work, signal connected clients to move new streams elsewhere, and allow in-flight work to finish within a bounded period. A forceful close after the grace period can still fail long-running requests, so the shutdown window must reflect actual request durations and client retry behavior.

Operational decision

A receipt API runs Go's standard HTTP server with HTTP/2 enabled on its TLS listener. First mark the instance unready and wait for upstream route removal. Then call Server.Shutdown with a bounded context; the Go fragment performs that application-side step, while the deployment must allow enough termination time for it to finish. Start several concurrent requests on one HTTP/2 connection, terminate the instance, and check that accepted requests complete while new requests move to other replicas. Repeat with a request longer than the shutdown budget and record its failure and retry outcome. Confirm the server process waits for Shutdown to return before exiting. Hijacked or upgraded connections need their own close path; the standard shutdown method does not wait for them. Inspect client connection reuse and error rates during the rollout, not only Pod readiness.

go
package main

import (
    "context"
    "net/http"
    "time"
)

func stopReceiptServer(receiptServer *http.Server) error {
    shutdownContext, cancelShutdown := context.WithTimeout(context.Background(), 9*time.Second)
    defer cancelShutdown()
    return receiptServer.Shutdown(shutdownContext)
}

Cost and verification

A nine-second grace window is an example, not a universal setting. Longer windows preserve slow requests but keep old replicas consuming capacity during a rollout; shorter windows cut them off. Compare the longest accepted request and the load balancer's route-removal time with the orchestrator's termination allowance. The shutdown call is cheap; the occupied connections and overlap replicas are the real costs. A zero error count on the server can conceal a client-side reset, so verify from both ends.

Common Mistakes

  • Do not equate readiness failure with closure of existing HTTP/2 streams.
  • Do not exit the process while graceful shutdown is still waiting.
  • Do not assume upgraded or hijacked connections are covered by Server.Shutdown.

Connected lessons

Practice and check

Advanced follow-up

devops
operations
Storage details