Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Kafka consumer offset recovery: preview every reset before changing a group

Last updated: 5 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

A Kafka consumer group commits positions for its assigned partitions. Processing and committing are separate operations, so a crash can cause an already applied event to be delivered again, or a premature commit can skip an unapplied event. An operator-initiated reset moves the group's future read position within the currently retained log. It does not recreate expired records or reverse a database transaction. Recovery therefore requires a per-partition plan, an external-effect reconciliation strategy, and evidence of the exact old and proposed positions. A reset should be previewed while group members are stopped so a running consumer cannot race the change.

Operational decision

A ledger consumer stops at offset 840 in partition two after malformed record 839. First capture the group state and output-ledger checkpoint. Fix the decoder or quarantine the poison event with its payload and reason; do not jump to latest. In a disposable rehearsal, reset only the affected topic and partition range to the chosen available offset, preview the new positions, and then execute after approval within the incident procedure. Restart a single consumer cohort and compare processed event IDs with ledger rows. If the record was already applied before the crash, its unique event ID prevents a second effect. Keep the old offsets and the CLI output in the incident evidence.

Output
Recovery record
Group: settlement-ledger-v2
Topic/partition: settlement-events/2
Old committed offset: 840
Proposed next offset: 839
Retained earliest offset: 711
Action: preview reset, execute while group inactive, replay with event-ID ledger
Abort if retained earliest offset exceeds required replay start

Cost and verification

Replaying R retained records has O(R) application work and can load databases or external APIs much harder than normal consumption. Throttle the recovery cohort and watch output latency, duplicate rejection, partition lag, and broker fetch load. A reset to earliest may multiply this cost across every partition; scope the change narrowly. If the needed offset is gone, restore a consistent snapshot plus surviving tail rather than pretending the skipped events were processed. A dead-letter queue is useful only when its ownership, repair, and replay path are tested.

Common Mistakes

  • Do not run an offset reset against an active group.
  • Do not use latest as a convenient way to clear lag.
  • Do not assume a reset rolls back external side effects.

Connected lessons

Practice and check

Redis operating follow-up

devops
kafka
stream-operations
Storage details