Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Kafka retention: size the replay window against the longest recovery path

Last updated: 2 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

A Kafka topic with delete retention removes old log segments according to time and size limits. Consumer offsets do not pin those records indefinitely. A consumer that remains offline past the available start offset cannot recover the missing interval by resetting its group. Retention therefore belongs in the service recovery contract, not just the broker storage configuration. Capacity calculations must account for ingress bytes, replication, compaction or segment overhead, and recovery traffic. A time target alone can be defeated by a smaller size cap; a size cap alone can be exhausted during a traffic burst.

Operational decision

For a settlement stream receiving an expected 47 GiB per day, a four-day replay requirement implies 188 GiB of logical new records before replication, overhead, safety margin, and burst allowance. Three replicas create substantially more physical storage demand. A two-day consumer outage still leaves only two days of headroom if traffic follows the forecast. Set alerts on the oldest available offset, consumer lag in time, and disk growth, then rehearse restoration from a snapshot when the log window is insufficient. When increasing retention, verify available broker capacity and the time needed to rebalance data. When reducing it, notify every replay-dependent consumer and preserve an alternate archive first.

Output
Expected ingress: 47 GiB/day
Required replay horizon: 4 days
Logical new-record floor: 47 x 4 = 188 GiB
Replica count: 3
Physical new-record floor before overhead: 564 GiB
Add burst, segments, reassignments, and disk safety margin

Cost and verification

Longer retention buys more recovery time but costs disk, replica transfer, and potentially longer catch-up after a broker loss. A small segment can improve deletion granularity while increasing file-management overhead; an active segment is not removed merely because an individual record is old. Compare the oldest retained record with the oldest needed restore point, not just a configured duration. A group lag metric of zero can also mislead if an operator reset the group to latest and skipped unrecovered events. Prove record continuity with source sequence or business IDs.

Common Mistakes

  • Do not assume committed offsets preserve records past topic retention.
  • Do not size only the leader's disk while ignoring replicas and rebalance headroom.
  • Do not reset to latest to hide a consumer that fell beyond the retained window.

Connected lessons

Practice and check

devops
kafka
stream-operations
Storage details