A systemd service can restart after selected exit failures. RestartSec delays the next attempt; StartLimitIntervalSec and StartLimitBurst constrain the number of starts in a window. When the limit is hit, later starts can be rejected until the condition is cleared or the interval passes. The correct values depend on how long the service needs to initialize and whether repeated starts damage a dependency. Restart=on-failure does not normally restart a service intentionally stopped by an operator. A crash-loop alarm should report the first meaningful failure and the current limit state rather than burying the cause under dozens of identical startup lines.
systemd restarts: bound crash loops and preserve evidence before resetting failures
Operational decision
A report renderer crashes immediately because a credential path is missing. Its unit permits four starts in twelve minutes with a 31-second restart delay. The fourth failure halts automated retries and pages the service owner. The operator captures the first exception, unit properties, credential mount state, and release ID. After restoring the credential, they test one explicit start and the report endpoint; only then do they reset the failed state if the start limit needs clearing. Increasing the burst count to 400 would spend CPU and flood logs without fixing the missing file. The same unit is tested for a recoverable single crash, a persistent configuration error, and a requested stop.
[Unit]
StartLimitIntervalSec=12min
StartLimitBurst=4
[Service]
Restart=on-failure
RestartSec=31s
Incident evidence: first error, last exit, restart count, release IDCost and verification
A longer delay slows recovery from a transient crash; a shorter delay can amplify a dependency outage and consume host resources. Rate limits also affect manual starts, so the runbook must say when a reset is justified. Monitor exit cause and failed-unit state separately. A process that catches every fatal error and keeps running may never trigger restart logic, while a process that exits zero after a failed job may be classified as successful. Verify the application's exit status contract alongside systemd's restart configuration.
Common Mistakes
- Do not increase the restart burst before diagnosing a deterministic failure.
- Do not reset failure state without retaining the first useful log and exit status.
- Do not assume an operator stop is equivalent to a process crash.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Linux service diagnostics: process, socket, and journal
- Alert design: page on impact and include a first action
- Credential incident response: revoke access before rebuilding trust
- systemd dependencies: separate unit ordering from application readiness
- systemd restarts: bound crash loops and preserve evidence before resetting failures
