Run this exercise in disposable Redis environments with synthetic reservation, catalog, and shipment data. Keep an external expected-state ledger before each fault. Do not use production credentials or assume one deployment topology supports every mode: a Sentinel replica drill and a Cluster slot migration need different isolated setups. The result is evidence about the chosen recovery contract, not a universal durability claim.
Redis recovery boundaries project: persistence, failover, resharding, and replay
Test persistence and memory
Write 147 numbered reservation keys, record the last acknowledged ID, and interrupt an instance during a background persistence operation. Restore from copied files and compare recovered IDs with receipts. Repeat a catalog-cache load at 134 percent of the forecast working set, observing evictions, write errors, resident memory, and source-database pressure. Keep authoritative reservations outside the evictable cache. The test fails if a memory policy silently removes a business token or if the backing database crosses its connection budget during refill.
Acceptance ledger
Persistence: report highest contiguous restored ID of 147
Cache overload: observe eviction and backing-store p99
Sentinel failover: list acknowledged IDs absent after promotion
Cluster move: redirect-aware clients reconcile final counters
Stream crash: one effect despite pending redelivery
Notification disconnect: 83 expirations reconciled from databaseInject failover and routing changes
Pause replication, acknowledge numbered writes on the primary, then fail it and compare the promoted replica against client receipts. Rebuild missing cache state from an authoritative ledger. In a separate Cluster test, move a slot while 29 clients update account-scoped counters; record MOVED and ASK handling and reconcile final values. A client error rate of zero is not enough if redirected retries apply an external effect twice.
Recover pending work and missed hints
Commit a shipment projection effect, crash before stream acknowledgment, claim the idle pending entry, and show that the unique event receipt prevents a duplicate projection. Pause a second group and test the proposed trim boundary against its oldest required entry. Disconnect a keyspace-notification subscriber while 83 reservation keys expire. Run a bounded database reconciliation and prove all 83 reach their intended state. Retain screenshots or command output for each fault, the expected ledger, and the measured recovery time.
Common Mistakes
- Do not infer durability from a successful write response alone.
- Do not declare failover complete before clients and state reconcile.
- Do not use a notification feed as the only record of an expiry-driven effect.
Connected lessons
- Redis persistence: state the recoverable write window before choosing AOF or snapshots
- Redis maxmemory: choose eviction behavior by the meaning of each key
- Redis Sentinel failover: reconcile acknowledged writes across a new primary
- Redis Cluster resharding: require redirect-aware clients and compatible key placement
- Redis Streams recovery: reconcile pending entries before trimming the log
- Redis keyspace notifications: keep correctness outside an ephemeral event channel
- DevOps projects
