A service outage on a Linux host can be caused by a dead process, a listener bound to the wrong address, a failed dependency, or a route that never reached the host. Diagnose those layers in order. A process listed as running is weaker evidence than a successful user request; process state only says the operating system has not terminated it.
Linux service diagnostics: process, socket, and journal
Operational decision
For a report-renderer service managed by systemd, inspect the unit state, recent journal entries, and listening sockets before restarting anything. Compare the process start time with the first bad request. If the service listens only on loopback while the proxy expects another interface, a restart may not change the fault. If systemd has repeatedly restarted the unit, capture the exit code and the first failure before changing the restart policy. Run the commands under an account with appropriate read access, then verify a request through the same route users take. A local curl success does not test DNS, TLS, or the edge proxy. Record the exact command output and timestamps in the incident timeline; a temporary fix without the original failure evidence makes later repair harder.
systemctl status report-renderer --no-pager
journalctl -u report-renderer --since '20 minutes ago' --no-pager
ss -ltnp | grep ':8147'
curl --fail --silent --show-error 127.0.0.1:8147/health/readyCost and verification
Journal retention and host logging consume disk; size and rotate them without deleting the evidence needed for incident review. A restart is fast when a process is wedged, yet repeated restarts can amplify load or hide a dependency fault. The local health request in the sample is one layer of evidence, not a production availability test. Check the edge path and a user-facing operation before declaring recovery. Limit privileged diagnostic access and avoid placing secret values in collected logs.
Common Mistakes
- Do not restart before recording the first error and exit state.
- Do not equate a listening socket with a healthy user operation.
- Do not treat a localhost check as proof that the public route works.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- DNS, TLS, and reverse-proxy failure boundaries
- Observability: join metrics, logs, and traces
- Incident response: contain impact, then learn
Advanced follow-up
- File descriptor exhaustion: find the leak before raising the limit
- Inode exhaustion: diagnose a full filesystem with free bytes
systemd operating follow-up
- systemd dependencies: separate unit ordering from application readiness
- systemd restarts: bound crash loops and preserve evidence before resetting failures
