Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Retries and timeouts: bound the cost of a failed request

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A timeout limits how long a caller waits for one operation. A retry attempts the operation again after failure; backoff with jitter spreads those attempts so clients do not hit a recovering dependency in lockstep. Multiple layers of retries can multiply downstream traffic. A transport timeout also cannot tell the caller whether a server committed the write before the connection broke.

Operational decision

A document-signing API calls a remote signer. Give the incoming operation a 2.4-second deadline, reserve time for serialization and response, and allow at most two remote attempts for errors classified as transient. Use a stable idempotency key for the same logical signing request so an uncertain first attempt does not create a second document. The pseudocode is a policy contract, not a library configuration; implement cancellation, error classification, and key persistence in the actual client. Do not retry validation failures or an authorization denial. For a rate limit, honor a trustworthy retry hint only if it fits within the remaining deadline. Observe attempt count, final failure rate, and work performed by the signer; a lower client error rate can hide a dependency being overwhelmed.

Output
Signing request policy
End-to-end deadline: 2400 ms
Remote attempts: at most 2 total
Attempt timeout: at most remaining deadline
Retry only: selected transport failure, 429, transient 5xx
Backoff: bounded exponential delay with jitter
Write identity: stable idempotency key per document request

Cost and verification

Two attempts can double load on a sick service. Nested client, proxy, and job retries can multiply that again. Tighter deadlines free resources sooner but may reject work that would have completed; choose them from measured latency and user tolerance. Store idempotency results long enough for the retry horizon and return the original result when the key repeats. A circuit breaker can reject work quickly during a sustained outage but needs its own recovery test.

Common Mistakes

  • Do not retry every error code.
  • Do not give each retry a fresh full request deadline.
  • Do not treat a timeout as proof that a write did not happen.

Connected lessons

Advanced follow-up

Advanced follow-up

Serverless operating-boundary follow-up

Redis operating follow-up

devops
resilience
Storage details