Kafka orders records within one partition, not across an entire topic. Producers commonly route a stable key to one partition so updates for one entity preserve order. A key chosen only for semantic convenience can create a hot partition: one large tenant or status value receives most of the traffic while other partitions sit idle. Adding consumers cannot split work inside a single assigned partition. Increasing partition count may change key-to-partition mapping for later records, so records for one entity can cross old and new partitions during a migration. Treat partition count and key format as a versioned data contract.
Kafka partition keys: balance throughput without breaking per-entity order
Operational decision
An invoice stream initially uses tenant ID as its key. One tenant produces 61 percent of writes, and its partition saturates while seven others remain underutilized. The service considers invoice ID as the new key because each invoice needs ordered state transitions, while cross-invoice order has no business meaning. Before switching, it checks that no consumer depends on tenant-wide order, dual-writes a shadow stream, and compares per-invoice state after replay. It assigns a new topic or deliberate migration boundary instead of simply expanding partitions under existing consumers. Scale-out estimates use the hottest partition rate, not the topic average.
Current key: tenant ID; hottest tenant: 61% of writes
Required order: all updates for one invoice
Candidate key: invoice ID
Migration: new topic + shadow comparison + bounded cutover
Capacity check: max partition ingress and consumer processing rateCost and verification
More partitions increase broker metadata, open files, replication work, rebalance time, and monitoring cardinality. A finer key spreads load only if traffic is actually distributed across those keys; a single invoice receiving all writes remains hot. Measure producer bytes and consumer processing time per partition, plus key distribution, before changing the contract. During migration, compare sequence numbers and final entity state, and define how late records on the old topic are drained. Avoid claiming exact per-key continuity if old and new topics process concurrently without a fencing or merge rule.
Common Mistakes
- Do not assume adding consumers increases parallelism beyond active partition count.
- Do not change partition count without checking key routing and ordering consequences.
- Do not use a low topic-wide average to dismiss one saturated partition.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Consumer rebalances: preserve ordering and effect ownership
- Queue-age scaling: target completion time, not only queue length
- Event schema evolution: release consumers before new event shapes
- Kafka write durability: align acknowledgments with the in-sync replica floor
- Kafka producer retries: preserve partition order without claiming end-to-end exactly-once
