An idle timeout closes a connection after a period without qualifying traffic. A client pool may still believe the socket is reusable when an intermediate proxy has closed it. The next request then sees a reset or an ambiguous failure. Idle timeout, maximum connection age, request timeout, and graceful drain are different controls; configuring one number across every hop can worsen throughput or still leave a race.
Proxy idle timers: stop reused sockets from failing the next request
Operational decision
A checkout client holds a pooled connection to a receipt gateway. The gateway closes idle backend sockets sooner than the client retires them, and the first request after a quiet period fails. Record each hop's idle timer, the client pool's idle eviction, maximum connection lifetime, protocol version, and retry policy. In an isolated test, send a request, wait just below and just above the suspected threshold, then send another request while logging the peer socket and any reset. Set the client's idle retirement earlier than the network hop it can observe, with a margin for scheduling and clock variance; test the exact client library because some validate a pooled socket before reuse. For a streaming response, confirm whether the proxy's timer is reset by application bytes or only by protocol traffic. Retry an idempotent read if its contract permits it, but reconcile an ambiguous mutation by operation ID before replay.
Receipt client pool contract
Gateway idle close: measured from gateway configuration
Client idle eviction: shorter than gateway close, with margin
Request deadline: independent of socket idle lifetime
Mutating retry: reuse operation key and reconcile prior result
Verification: first request after quiet period succeedsCost and verification
Short pool lifetimes increase DNS lookups, TCP and TLS handshakes, CPU, and ephemeral port churn. Long lifetimes improve reuse but extend old-endpoint exposure and increase idle-close races. A proxy may have separate client-facing and backend-facing timers, so tune the failing hop rather than the most visible one. Measure handshake rate, first-after-idle errors, request latency, and reset direction. Verify from both client and proxy logs; a backend success counter can miss a response lost at the edge.
Common Mistakes
- Do not call every timeout a request deadline.
- Do not retry an ambiguous payment with a new operation key.
- Do not assume one load-balancer timer covers both sides of the proxy.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Client connection pools: retire old endpoints after a DNS or rollout change
- Retries and timeouts: bound the cost of a failed request
- Idempotency keys: reconcile an accepted write before repeating it
- Connection draining: let in-flight work finish while new traffic moves
