Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Kafka write durability: align acknowledgments with the in-sync replica floor

Last updated: 5 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

A Kafka partition has a leader and replica followers. The in-sync replica set, or ISR, contains replicas close enough to the leader to qualify for an acknowledged write. The producer's acknowledgments setting and the topic's minimum ISR jointly decide whether a write succeeds during a broker fault. Replication factor describes assigned copies, not how many are currently in sync. This distinction matters during maintenance: a three-copy topic can temporarily have one eligible replica. An application that treats a successful send as a durable business event must record the exact producer and topic contract, then test the failure behavior.

Operational decision

For a payment-event topic, assign three replicas, require two in-sync replicas, and use producer acknowledgments of all. If one broker is unavailable while two replicas remain in sync, writes may continue; if the ISR falls below two, writes fail instead of relaxing the policy silently. When all ISR members are unavailable, electing an out-of-sync replica can restore availability but may discard acknowledged history. Keep that choice explicit in the incident runbook. Producers must surface send failures to callers and retain an idempotent business request for retry. Measure ISR shrinkage, under-replicated partitions, producer errors, and the time taken to restore a replica before approving planned broker maintenance.

Output
Topic: payment-events
Replication factor: 3
min.insync.replicas: 2
Producer acks: all
ISR size 3: accept after all 3 ISR replicas acknowledge
ISR size 2: accept after both acknowledge
ISR size 1: reject the write

Cost and verification

Waiting for more replicas raises write latency and network work; a strict ISR floor can turn a replica outage into a write outage. That is an intentional trade. Capacity planning must leave disk, network, and broker headroom for follower catch-up; otherwise the system can remain below the write floor long after the initial fault is repaired. No setting promises zero loss under every correlated failure, disk fault, or incorrect operational intervention. Verify the actual client result, accepted offset, and post-failover record visibility in a disposable cluster rather than relying on a configuration screenshot.

Common Mistakes

  • Do not equate replication factor with current ISR size.
  • Do not describe acks=all as an absolute zero-loss guarantee.
  • Do not enable unclean leader election merely to clear an availability alert without accepting its loss risk.

Connected lessons

Practice and check

RabbitMQ operating follow-up

devops
kafka
stream-operations
Storage details