A timeout means the caller stopped waiting; it does not prove the server did not commit the operation. Retrying blindly can duplicate work and amplify an outage. Give a request an end-to-end deadline, allow only a small number of attempts inside it, and retry only errors that are transient and safe for that operation. A GET is generally easier to retry than a mutation with unknown commit outcome; a mutation needs an idempotency contract or a status lookup. Backoff with jitter reduces synchronized repeat traffic. Honor an explicit server retry delay when applicable, while still respecting the caller's own deadline.
Request Deadlines, Retries, and Backoff
Working case
The case service sends an update to an image-analysis dependency. The connection drops after the dependency accepts the job, and the local handler cannot tell whether it started. A client that retries immediately may create duplicate analysis work. Give the operation a stable idempotency key and query its status where supported. If a transient 503 occurs before acceptance, wait within a bounded retry budget using varied delays. The reviewer should see whether the action is still pending, definitively failed, or safe to retry; a generic spinner after the deadline is not a contract.
Implementation boundary
function mayRetryRequest({ status, attempt, remainingMs, repeatSafe }) {
if (!repeatSafe || attempt >= 3 || remainingMs < 250) return false;
return status === 429 || status === 502 || status === 503 || status === 504;
}
console.log(mayRetryRequest({ status: 503, attempt: 2, remainingMs: 900, repeatSafe: true }));
// Output: trueThis predicate decides eligibility, not the delay. A production caller should attach an absolute deadline to all attempts, cap each connection and response wait, and compute bounded backoff with jitter. A server-provided retry delay is advice, not permission to exceed the user's remaining time. Decide which HTTP outcomes are retryable for the specific operation; 429 and 503 may be candidates, whereas a validation error or stale If-Match condition requires user or client correction. Do not stack independent retry loops in browser, API gateway, and downstream client without a total attempt budget.
Cost and boundaries
One attempt costs one request; three attempts can triple work during a partial outage. Backoff lowers synchronized pressure but increases latency for the affected user. A deadline bounds that exposure and makes failure visible, while a too-short deadline can abandon work that would have completed. Measure success after retry, extra attempt count, dependency saturation, and end-to-end latency. An idempotency record or operation status endpoint adds storage and cleanup cost but is often cheaper than duplicate side effects.
Failure trace
A client retries every 500 response three times, while the gateway and downstream SDK also retry. One user action fans out into many calls just as the dependency is failing. Another client retries a timed-out POST without a stable key, creating two analysis jobs. Define one retry owner, cap total attempts and time, and make the mutation repeat-safe before automatic retry. Test a timeout after commit, an immediate 503, a rate limit with retry advice, and a deadline that expires during backoff.
Verification
- A timeout after commit does not create a duplicate job on retry.
- Transient retries stop at the attempt and overall-deadline limits.
- Validation and stale-version failures are not retried as transient errors.
Practice drill
Instrument one case-analysis request with an idempotency key and a five-second overall deadline. Force a 503 before acceptance, a dropped response after acceptance, and a validation error. Record attempted calls, wait time, final user state, and downstream job count. Add jitter to concurrent callers and compare traffic bursts. Then enable a second retry layer on purpose to see amplification, remove it, and document the single owner of the retry budget.
Decision note
Retry only when the operation is safe to repeat, the failure is plausibly transient, and the caller still has time budget.
Common Mistakes
- Equating timeout with proof that no side effect occurred.
- Allowing every network layer to run its own retry loop.
- Ignoring a remaining deadline while waiting through backoff.
Connected lessons
API Mutation and Failure Contracts; Conditional Writes and Lost-Update Prevention; Accepted Operations and Status Resources; Structured API Errors and Recovery; Idempotent Write Requests and Lost Responses; Rate Limits and Request Budgets; Observability: connect user failure to a safe request trace.
Apply and check
Build Project: conflict-safe case API and review Web Development: API mutation contracts quiz.
Further connections
Proxy Body Budgets, Timeouts, and Retry Ownership.
Further connections
Angular HttpClient Streams and Stale Detail Requests.
