A cache stampede occurs when many callers miss the same key and independently request the same expensive source value. It is especially damaging when a popular entry expires across replicas at the same instant or when the cache restarts empty. A per-process singleflight group can collapse concurrent loads inside one process, but six application replicas may still make six source calls. A shared refresh lease, controlled stale serving, staggered expiry, or admission limit can bound the cross-process work. Each option has a failure case. A lease holder may crash; stale data may be forbidden for a private decision; randomizing expiry only spreads scheduled misses and cannot repair a sudden eviction.
Cache Stampede, Singleflight, and Hot-Key Budgets
Working case
The permit dashboard asks for a daily district 47 count. At 09:00, 900 browsers refresh while six API replicas see the same expired key. The aggregate query takes 430 milliseconds and scans several joined tables. Without coordination, hundreds of copies run and delay case writes. A process-local singleflight reduces each replica to one query, but six still hit the database. The service adds a short shared refresh lease with an expiry, serves a recent count for at most 47 seconds when permitted, and limits source refresh concurrency. Case approval and access checks continue to use the source because those decisions cannot use that stale count.
Implementation boundary
function sourceLoadsWithoutSharedLease(replicaCount, processSingleflightEnabled) {
return processSingleflightEnabled ? replicaCount : 900;
}
console.log(sourceLoadsWithoutSharedLease(6, true));
// Output: 6First identify the key and its actual source cost, request rate, acceptable age, and tenant scope. Collapse simultaneous loads within each instance using one promise per key, then decide whether cross-instance coordination is justified by measured source pressure. A refresh lease needs an expiry so a crashed holder does not freeze the value forever; writes from an old holder must be rejected if a later holder has a newer generation. Give waiting callers a deadline and a fallback: bounded stale count, source request under an admission budget, or an explicit temporary error. Jitter expiration for independent keys so scheduled refreshes do not synchronize. Warm only a measured hot set after restart; preloading every possible tenant key can be more expensive than normal misses. Keep a hard cap on concurrent source work.
Cost and boundaries
Singleflight uses O(K) in-flight coordination for K active keys in one process and retains a promise only until the load settles. A shared lease adds network operations and failure handling. Stale serving lowers source pressure but trades freshness for availability within a declared age. Jitter reduces synchronized expiry but does not prevent a cache flush stampede. A too-long lease leaves callers waiting behind a dead worker; a too-short lease permits duplicate refreshes before a slow query finishes. Measure misses per hot key, source calls per miss burst, lease wait, stale age, rejected refreshes, source p95, and user task completion. Do not hide all failures behind a stale response forever.
Failure trace
Expire a district count on all six replicas and send 900 simultaneous reads. Count source queries, not just cache hit rate after the burst. Crash the lease owner before it fills the value and verify another worker can take over after the lease expiry. Slow the source query beyond the lease duration; an old worker must not overwrite a newer generation. Flush the whole cache while a rollout doubles app replicas and observe that source concurrency stays under its budget. Ask for a private approval in the same test; it must not inherit the count's stale-data permission.
Verification
- Burst tests measure source calls per key, not only steady-state hits.
- A crashed or slow lease owner cannot pin or overwrite the result indefinitely.
- Stale serving is limited to routes with an explicit age budget.
Practice drill
Set a 430-millisecond count query, 900 concurrent readers, and six application replicas. Compare uncoordinated loading, process-local singleflight, and a shared refresh lease. Define a 47-second maximum stale age for the count, a shorter lease lifetime with safe renewal, and a 63-query aggregate source budget. Then kill the refresher, delay its old result, and let a replacement finish first. Record which result may be cached, how many source calls ran, and what response late readers receive.
Decision note
Collapse work at the scope where the stampede occurs, with a deadline and a source-load ceiling.
Common Mistakes
- Assuming process-local singleflight coordinates every replica.
- Using a refresh lock with no expiry or fencing rule.
- Serving stale permission or payment state because stale counts are allowed.
Related lessons
Application Cache Consistency and Capacity; Cache-Aside Fill Races and Version Guards; Cache Memory, Eviction, and Degraded Read Paths; Negative Cache Entries and Stale-Read Contracts; Load Tests and Capacity Budgets; Connection Pool Budgets and Queue Admission.
Apply and check
Build Project: permit cache recovery and consistency and review Web Development: application cache contracts quiz.
