Queue depth shows unfinished work but does not by itself describe how long the oldest item has waited. A worker scaler may include in-flight messages, which changes the meaning of its target. More consumers can shorten backlog only while the database, remote API, and broker can serve the added concurrency. Delayed messages, poison items, and long visibility leases can distort a single depth number.
Queue-age scaling: target completion time, not only queue length
Operational decision
A billing queue holds 2,400 receipts after a provider outage. Measure visible count, in-flight count, oldest eligible age, arrival rate, completion rate, and per-worker processing time. Set a temporary age objective and calculate the drain rate needed to meet it, then cap worker concurrency at the downstream payment API and database budgets. This policy sample makes the limit explicit; implement the equivalent controls in the actual scaler and worker. Test scale-from-zero by publishing one item, including the time until a worker is ready. Keep a small warm floor when that cold-start delay violates the age objective. During recovery, watch completed receipts and oldest age; a falling visible count caused only by long in-flight leases is not successful drainage. Reduce the worker ceiling if provider throttling or database wait time climbs. Replay dead-lettered work under its own rate budget after normal intake stabilizes.
Billing queue recovery gate
Oldest eligible age objective: under 6 minutes
Observe: visible, in-flight, delayed, completed/minute, failure/minute
Worker ceiling: min(queue demand, database slots, provider rate budget)
Scale-from-zero test: one receipt reaches durable completion
Rollback: lower worker ceiling if downstream wait or 429 rate risesCost and verification
A warm worker floor increases compute cost. Aggressive scaling can spend more while reducing completed throughput if every worker contends for the same dependency. Age is a service outcome; use depth and throughput to explain it. Estimate the drain window from net completion rate after new arrivals, not a single static queue count. Keep the lease and idempotency contract valid as replica count changes.
Common Mistakes
- Do not count in-flight messages as completed work.
- Do not choose a replica target without measuring processing time and downstream capacity.
- Do not assume scale-from-zero meets an age objective until the cold path is tested.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Queue consumers: acknowledgement, idempotency, and backlog
- Queue visibility leases: prevent overlapping workers on one message
- Database pool pressure: bound waiting before the database collapses
- HPA stabilization: prevent replica oscillation without masking demand
Practice and check
Kafka operating follow-up
- Kafka retention: size the replay window against the longest recovery path
- Kafka partition keys: balance throughput without breaking per-entity order
