Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Kafka stream recovery project: broker fault, replay, and state rebuild

Last updated: 2 Oct 202612 min read
project
AdvancedBy AITrove Editorial

Build a disposable three-broker stream for synthetic settlement and customer-preference events. Give the settlement topic three replicas and a two-replica write floor; put keyed customer state in a separate compacted topic. Record every generated event ID and expected final state outside the cluster before starting the fault drill. This project is a rehearsal, not a production change: isolate credentials, volumes, and network access so a reset or broker stop cannot affect live groups.

Establish the write boundary

Publish 240 uniquely identified settlement records with idempotent producer behavior and acknowledgments of all. Stop one follower and confirm writes continue while two replicas remain in sync. Stop a second eligible replica; confirm the producer receives an error and the caller retains its request for retry. Restore replicas and verify accepted IDs from the surviving log. Do not claim the exercise succeeded merely because the producer process exited zero: compare broker records, caller receipts, and the external event ledger.

Output
Synthetic acceptance
Settlement records requested: 240 unique IDs
Accepted IDs before fault: reconcile with broker log
ISR floor: 2 of 3 replicas
Below floor: write rejected, request retained for retry
State rebuild: deleted customer key absent
Offset reset: preview captured before execute; group inactive
Expired replay start: restore snapshot plus surviving tail

Inject replay faults

Lose a producer acknowledgment and resend the same batch; compare broker append IDs with the business event ledger. Pause a consumer, let retained segments advance in the disposable topic, and show why a reset cannot recover records older than the earliest offset. Restore a known snapshot and replay the surviving tail without repeating external settlement effects. Seed a customer key with two values and a tombstone, then bootstrap a fresh view while the cleaner is active; it must end with the key absent. Record bootstrap duration against the configured deletion-marker window.

Change the partition contract

Create a skewed stream where one tenant owns most writes. Measure per-partition ingress and consumer time; then shadow a new topic keyed by invoice ID. Compare per-invoice sequence and final state before declaring cutover. Finally stop the test consumer group, capture old and proposed offsets, preview a narrow reset, execute it, and start one cohort. Reconcile duplicates with unique business IDs. Archive the commands, observations, and rejected assumptions as the project evidence.

Common Mistakes

  • Do not run a reset while consumers are active.
  • Do not count rejected sends as accepted events.
  • Do not interpret a low topic average as proof that every partition has spare capacity.

Connected lessons

devops
project
Storage details