A Kafka partition has a leader and replica followers. The in-sync replica set, or ISR, contains replicas close enough to the leader to qualify for an acknowledged write. The producer's acknowledgments setting and the topic's minimum ISR jointly decide whether a write succeeds during a broker fault. Replication factor describes assigned copies, not how many are currently in sync. This distinction matters during maintenance: a three-copy topic can temporarily have one eligible replica. An application that treats a successful send as a durable business event must record the exact producer and topic contract, then test the failure behavior.
Kafka write durability: align acknowledgments with the in-sync replica floor
Operational decision
For a payment-event topic, assign three replicas, require two in-sync replicas, and use producer acknowledgments of all. If one broker is unavailable while two replicas remain in sync, writes may continue; if the ISR falls below two, writes fail instead of relaxing the policy silently. When all ISR members are unavailable, electing an out-of-sync replica can restore availability but may discard acknowledged history. Keep that choice explicit in the incident runbook. Producers must surface send failures to callers and retain an idempotent business request for retry. Measure ISR shrinkage, under-replicated partitions, producer errors, and the time taken to restore a replica before approving planned broker maintenance.
Topic: payment-events
Replication factor: 3
min.insync.replicas: 2
Producer acks: all
ISR size 3: accept after all 3 ISR replicas acknowledge
ISR size 2: accept after both acknowledge
ISR size 1: reject the writeCost and verification
Waiting for more replicas raises write latency and network work; a strict ISR floor can turn a replica outage into a write outage. That is an intentional trade. Capacity planning must leave disk, network, and broker headroom for follower catch-up; otherwise the system can remain below the write floor long after the initial fault is repaired. No setting promises zero loss under every correlated failure, disk fault, or incorrect operational intervention. Verify the actual client result, accepted offset, and post-failover record visibility in a disposable cluster rather than relying on a configuration screenshot.
Common Mistakes
- Do not equate replication factor with current ISR size.
- Do not describe acks=all as an absolute zero-loss guarantee.
- Do not enable unclean leader election merely to clear an availability alert without accepting its loss risk.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Queue consumers: acknowledgement, idempotency, and backlog
- Consumer rebalances: preserve ordering and effect ownership
- Recovery-time budgets: measure every stage of service return
- Kafka write durability: align acknowledgments with the in-sync replica floor
- Kafka producer retries: preserve partition order without claiming end-to-end exactly-once
Practice and check
- Kafka stream recovery project: broker fault, replay, and state rebuild
- Kafka operating decisions: durability, retention, and recovery quiz
