systemd-journald can keep records in volatile storage or on a persistent filesystem, depending on configuration and directory state. A service's stdout and stderr can enter the journal, but an application that logs a credential or personal payload leaks it there too. Size limits bound stored journal files, yet active files and disk pressure from other writers can make real usage differ from a simple configured number. A node restart may erase volatile logs that would explain the outage. A production incident workflow therefore checks storage mode, the available retention horizon, forwarding health, and the exact service and boot identifiers needed to query an earlier failure.
systemd journals: retain failure evidence without exhausting the host disk
Operational decision
A payment proxy enters a crash loop overnight. The operator needs the first error from the previous boot, not just the latest restart. On a disposable host, they set persistent journaling with a 768 MiB system budget, generate controlled log volume, reboot, and query records by unit and previous boot. They confirm the first failure remains available long enough for incident response and that log forwarding receives the same event ID. They then simulate a disk filling from another process and verify whether journal retention behaves as expected. Sensitive request bodies are redacted at the application boundary; shrinking journal retention is not a data-protection substitute.
[Journal]
Storage=persistent
SystemMaxUse=768M
Review: previous-boot records for payment-proxy.service
Forwarding: alert if delivery stalls
Log contract: event ID present, credential and payload absentCost and verification
More local history consumes disk and write bandwidth; less history weakens incident reconstruction when central forwarding fails. Active journal files may not be removed immediately to meet a vacuum target. Rate limits can also suppress repetitive events, making an absence of logs ambiguous during a storm. Monitor journal disk use, forwarding lag, and dropped-message indicators. Set the budget against the same filesystem's application data and recovery reserve. A unit that emits large stack traces on every 31-second restart can exhaust a small budget before anyone arrives to investigate.
Common Mistakes
- Do not assume previous-boot logs exist with volatile storage.
- Do not place secrets or full sensitive payloads in journal-bound output.
- Do not interpret a size limit as an exact immediate disk-usage ceiling.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Log pipelines: preserve incident evidence without ingesting secrets
- Incident response: contain impact, then learn
- File descriptor exhaustion: find the leak before raising the limit
- systemd dependencies: separate unit ordering from application readiness
- systemd restarts: bound crash loops and preserve evidence before resetting failures
