Linux connection tracking stores state for network flows used by firewall and address-translation paths. A full table can reject new flows even when the application has spare workers and the load balancer reports healthy. The count and maximum describe one host's tracking capacity, not the number of useful application requests. Timeouts, UDP traffic, and NAT behavior determine how long entries remain, so a burst can outlast the traffic that created it.
Connection tracking: separate host flow exhaustion from application limits
Operational decision
A receipt service starts timing out on new outbound ledger calls after a deployment doubles connection churn. On each affected Linux node, collect the read-only count and maximum below, kernel drop or insert-failure counters, active sockets, and per-destination traffic. Compare nodes carrying equal request volume; a single hot node can reach its limit first. Then determine whether the increase came from lost connection reuse, retry amplification, an egress proxy, or a scan. Restore pool reuse or pace retries before raising limits. If a larger table is justified, budget kernel memory and test it under the same burst and failure conditions. Check that new connections recover without deleting existing legitimate flows. A successful application readiness check can still miss this fault if it reuses one established connection.
cat /proc/sys/net/netfilter/nf_conntrack_count
cat /proc/sys/net/netfilter/nf_conntrack_max
ss -sCost and verification
Tracking more flows consumes memory and increases inspection work; lowering timeouts can interrupt legitimate idle protocols. A blanket maximum increase may hide a runaway client until the next, larger burst. Inspect the ratio of count to limit, failed inserts, memory pressure, and client-visible connection errors together. The two files exist on Linux systems with connection tracking enabled; container and namespace views may differ from the node that actually performs translation.
Common Mistakes
- Do not equate open application sockets with every tracked flow.
- Do not raise the limit without finding the flow source and memory cost.
- Do not count a green readiness probe as proof that new outbound flows can be allocated.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- NAT port pressure: find the shared outbound ceiling
- Retry amplification: assign one owner for each failed operation
- File descriptor exhaustion: find the leak before raising the limit
- Load-balancer health: a green socket is not a working transaction
