A dead-letter destination receives messages a queue could not process under its configured failure policy. It is a holding area for investigation, not proof that the event was harmless or will be retried automatically. Replaying a message can duplicate a completed side effect when the consumer committed work but failed before acknowledgement.
Dead-letter replay: recover failed messages without repeating their effects
Operational decision
A royalty worker rejects a malformed payout event after a bounded attempt count. Record its event ID, failure class, original enqueue time, and the consumer revision that failed; do not put personal data into an unrestricted dead-letter log. The policy sketch is not a broker declaration. Fix the parser or data contract, then select a small replay cohort by event ID and send it through the normal consumer with the original business idempotency key. Verify ledger effects and acknowledgement count before expanding the cohort. A permanently invalid event should move to a documented manual disposition, not spin in a retry loop. Keep a separate alarm for dead-letter growth and oldest age. The broker may provide different delivery guarantees for dead-letter transfer, so test failure of the transfer path itself.
Royalty replay gate
Original event ID: preserved
Failure class: transient, contract, or permanent
Attempt cap: 3 consumer deliveries
Replay cohort: named IDs, operator, start time
Effect guard: unique payout operation key
Stop gate: duplicate ledger write or rising failure ageCost and verification
Dead-letter storage and investigation add cost, but infinite retries can saturate the main queue and hide useful work. Replay throughput should be limited to protect the recovered dependency. Message age matters: the business action may no longer be valid even when the parser now accepts it. Observe the original effect boundary, not only broker acknowledgements. Retain enough metadata to audit a replay while applying the same data-retention controls as the source queue.
Common Mistakes
- Do not replay every failed event at once after a fix.
- Do not assign a new idempotency key to the same logical payout.
- Do not equate an empty dead-letter queue with no lost work.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Queue consumers: acknowledgement, idempotency, and backlog
- Retries and timeouts: bound the cost of a failed request
- Alert design: page on impact and include a first action
- API compatibility windows: release consumers and producers safely
Practice and check
Object storage follow-up
Kafka operating follow-up
- Kafka consumer offset recovery: preview every reset before changing a group
- Kafka stream recovery project: broker fault, replay, and state rebuild
RabbitMQ operating follow-up
- RabbitMQ dead lettering: make poison-message transfer observable and recoverable
- RabbitMQ delivery recovery project: confirms, quorum, and queue cutover
