A process uses file descriptors for files, sockets, pipes, and other kernel objects. When it reaches its per-process open-file limit, new opens can fail even when the host has free memory and disk space. Raising the limit may postpone the failure, but a leaking connection or unclosed file can consume the new allowance too. A gateway with many keep-alive connections also needs a deliberate capacity budget rather than an arbitrary maximum.
File descriptor exhaustion: find the leak before raising the limit
Operational decision
A document API reports intermittent connection failures after several hours of traffic. On a Linux host using systemd, read its main PID, process limits, and descriptor count with the shell fragment. The commands are read-only; use a namespace-aware equivalent when the service runs in a container. Sample the count over time and classify descriptors through a bounded inspection of /proc or a diagnostic tool. Compare the trend with active clients, connection pool size, and deployment age. Reproduce the workload in a disposable environment and close the request body stream and upstream response in every path, including timeout and exception paths. Then confirm the descriptor count returns toward baseline after a traffic burst. If the service legitimately needs more concurrent connections, change the service and host limits together, account for memory per connection, and test the load balancer's own connection ceiling.
process_id=$(systemctl show -p MainPID --value document-api.service)
cat "/proc/$process_id/limits"
find "/proc/$process_id/fd" -mindepth 1 -maxdepth 1 | wc -lCost and verification
Each open connection consumes kernel and often application memory. A higher file limit can improve legitimate concurrency but can also let a leak damage the host more deeply before failing. Frequent full descriptor enumeration has a cost on busy processes; use it for diagnosis, not a high-frequency metric scrape. Track open descriptors, accept failures, and user request errors together. Restarting the process clears its descriptors temporarily but does not repair the leak.
Common Mistakes
- Do not raise the open-file limit as the only fix for a rising leak.
- Do not count disk free space as evidence that descriptors are available.
- Do not omit timeout and exception paths when checking resource closure.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Linux service diagnostics: process, socket, and journal
- Graceful Pod shutdown: stop accepting work before exit
- Database pool pressure: bound waiting before the database collapses
- Observability: join metrics, logs, and traces
