Overload shedding rejects work a service cannot finish within its deadline. It is different from a crash or silent queue growth: the caller receives an explicit response, and the service protects capacity for work already admitted. Autoscaling may add capacity, but it reacts after measurement and cannot help when a shared database, signer, or queue is the bottleneck.
Overload shedding: refuse excess work before latency collapses
Operational decision
A document-signing API has 47 workers and a measured capacity of 82 concurrent signing requests. Admit work only while the bounded queue remains below its deadline-based limit; reject excess requests with a retryable status and a short, bounded retry hint. The text policy is an application contract, not a proxy configuration. Choose which requests retain priority: a completion callback may matter more than a new preview request. Test with a burst that exceeds capacity and verify that admitted operations still complete within the objective. Measure rejected requests separately from failures, because hiding the shed rate makes the dashboard look healthy while users are refused. Clients must use jitter and a total deadline; identical immediate retries can make the overload worse.
Document signing admission policy
Workers: 47
Measured safe in-flight work: 82 requests
Queue admission: only if predicted finish fits deadline
Over-capacity result: explicit retryable response
Client contract: bounded attempts, jitter, stable operation key
Measure: admitted latency, shed rate, downstream saturationCost and verification
A small queue improves latency under overload but rejects more requests; a large queue increases apparent acceptance while many requests expire before service. The correct limit depends on service time distribution, not only worker count. Reserved capacity for high-priority operations can protect core workflows but needs fairness rules to prevent starvation. Test the behavior when telemetry or the admission counter fails, and decide whether to fail closed or use a conservative static limit.
Common Mistakes
- Do not use an unbounded queue to disguise overload.
- Do not let clients retry shed work immediately in sync.
- Do not measure only admitted requests when reporting availability.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Retries and timeouts: bound the cost of a failed request
- Horizontal autoscaling: choose a signal tied to demand
- Capacity and load tests: identify the next bottleneck
- SLOs and error budgets: turn reliability into a decision
Practice and check
Advanced follow-up
- Cache stampedes: bound origin work when popular keys expire
- Edge rate limits: reject abuse without penalizing shared networks
