Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Redis recovery boundaries project: persistence, failover, resharding, and replay

Last updated: 5 Oct 202612 min read
project
AdvancedBy AITrove Editorial

Run this exercise in disposable Redis environments with synthetic reservation, catalog, and shipment data. Keep an external expected-state ledger before each fault. Do not use production credentials or assume one deployment topology supports every mode: a Sentinel replica drill and a Cluster slot migration need different isolated setups. The result is evidence about the chosen recovery contract, not a universal durability claim.

Test persistence and memory

Write 147 numbered reservation keys, record the last acknowledged ID, and interrupt an instance during a background persistence operation. Restore from copied files and compare recovered IDs with receipts. Repeat a catalog-cache load at 134 percent of the forecast working set, observing evictions, write errors, resident memory, and source-database pressure. Keep authoritative reservations outside the evictable cache. The test fails if a memory policy silently removes a business token or if the backing database crosses its connection budget during refill.

Output
Acceptance ledger
Persistence: report highest contiguous restored ID of 147
Cache overload: observe eviction and backing-store p99
Sentinel failover: list acknowledged IDs absent after promotion
Cluster move: redirect-aware clients reconcile final counters
Stream crash: one effect despite pending redelivery
Notification disconnect: 83 expirations reconciled from database

Inject failover and routing changes

Pause replication, acknowledge numbered writes on the primary, then fail it and compare the promoted replica against client receipts. Rebuild missing cache state from an authoritative ledger. In a separate Cluster test, move a slot while 29 clients update account-scoped counters; record MOVED and ASK handling and reconcile final values. A client error rate of zero is not enough if redirected retries apply an external effect twice.

Recover pending work and missed hints

Commit a shipment projection effect, crash before stream acknowledgment, claim the idle pending entry, and show that the unique event receipt prevents a duplicate projection. Pause a second group and test the proposed trim boundary against its oldest required entry. Disconnect a keyspace-notification subscriber while 83 reservation keys expire. Run a bounded database reconciliation and prove all 83 reach their intended state. Retain screenshots or command output for each fault, the expected ledger, and the measured recovery time.

Common Mistakes

  • Do not infer durability from a successful write response alone.
  • Do not declare failover complete before clients and state reconcile.
  • Do not use a notification feed as the only record of an expiry-driven effect.

Connected lessons

devops
project
Storage details