Redis Streams consumer groups track delivered but unacknowledged entries in a pending entries list. A worker that crashes after a delivery leaves an entry pending; another worker can inspect and claim it after an appropriate idle interval. Claiming transfers responsibility, not proof that the previous worker failed before its external effect. A unique business event ID must guard the downstream effect before acknowledgment. Stream trimming is a separate storage action: removing an entry before every required consumer group has finished can leave a pending reference whose payload is unavailable. Retention must be sized against the slowest recovery path, not only the fastest group.
Redis Streams recovery: reconcile pending entries before trimming the log
Operational decision
A shipment projector reads stream shipments through group tracking-v2. Worker A writes projection event sh-7284, then crashes before XACK. Worker B waits beyond the measured normal processing interval, claims the idle entry, checks the projection's unique event receipt, and acknowledges without writing a second projection. The operator compares pending age, entry payload availability, last delivered ID, and projection count. A second analytics group has a longer replay window, so trim limits are chosen for that group too. If a required entry has already been trimmed, the team restores a snapshot and surviving stream tail rather than treating an empty payload as successful completion.
Stream: shipments
Group: tracking-v2
Event ID: sh-7284
Worker A: effect committed, acknowledgment missing
Worker B: claim idle entry, check unique receipt, then XACK
Trim gate: oldest required group recovery point remains availableCost and verification
Scanning and claiming pending work costs Redis CPU and can flood a downstream system if a large backlog is released at once. Set the idle threshold above a realistic high percentile of processing time, then throttle claims. A threshold that is too short can make two healthy workers contend for the same effect. Track pending count and oldest pending age per group, not just stream length. Treat a trim policy as data deletion. Before shortening it, verify snapshots, lagging groups, and any consumer that is intentionally paused for maintenance.
Common Mistakes
- Do not interpret a claim as proof the previous worker did no work.
- Do not trim a stream solely by its fastest consumer group.
- Do not acknowledge an unavailable payload as if it were processed.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Dead-letter replay: recover failed messages without repeating their effects
- Queue consumers: acknowledgement, idempotency, and backlog
- Kafka consumer offset recovery: preview every reset before changing a group
- Redis persistence: state the recoverable write window before choosing AOF or snapshots
- Redis maxmemory: choose eviction behavior by the meaning of each key
