A partitioned event log orders records within each partition, not across an entire topic. Consumers in a group divide partition ownership; a membership change can reassign a partition. A consumer that commits its offset before its business effect is durable can skip work after a crash. A consumer that performs the effect and crashes before committing may process that record again after rebalance.
Consumer rebalances: preserve ordering and effect ownership
Operational decision
A payout stream uses account ID as the partition key so updates for one account retain partition order. Process each record under its original operation key, commit the ledger effect, then commit the corresponding offset. The text block is a handoff contract, not broker configuration. During a rolling update, stop fetching new records, finish or abandon in-flight work within the shutdown deadline, and let the next owner retry uncommitted records. Verify duplicate delivery by killing a worker after the ledger commit but before offset commit. Adding partitions may change the hash-to-partition mapping for future records, so test account ordering across the expansion boundary before assuming a stable key gives the same partition forever. Record consumer lag by partition and the oldest unprocessed event age; an aggregate lag can hide one hot account.
Payout partition handoff
Key: account ID
Effect key: original payout operation ID
On record: validate, persist ledger effect, then commit offset
On revoke: stop fetch; finish bounded in-flight work
On restart: retry uncommitted record under same effect key
Check: one ledger effect after forced rebalanceCost and verification
More partitions can increase parallelism but also metadata, open fetches, and rebalance work. A single hot key remains serialized on one partition even when worker count rises. Committing offsets in every record can add broker traffic; batching commits reduces overhead but increases the retry window. Idempotency at the ledger boundary is therefore necessary even with careful offset order. Test a rebalance under load, not only a clean shutdown.
Common Mistakes
- Do not claim global event order across partitions.
- Do not commit an offset before the business effect is durable.
- Do not change partition count without checking key ordering assumptions.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Queue consumers: acknowledgement, idempotency, and backlog
- Dead-letter replay: recover failed messages without repeating their effects
- Event schema evolution: release consumers before new event shapes
- Graceful Pod shutdown: stop accepting work before exit
