Skip to content
AITroveRead. Build. Understand.
Make this comfortable

systemd service fault drill: readiness, restart, timer, and socket recovery

Last updated: 2 Oct 202612 min read
project
AdvancedBy AITrove Editorial

Use a disposable Linux VM with systemd and a synthetic service called settlement-api. Version-control its service, timer, and socket unit definitions. Before applying them, verify unit syntax with the installed systemd tooling and inspect the effective settings. Keep one ledger of synthetic request IDs and completed effects outside the service process. The exercise tests lifecycle behavior; do not reuse a production credential or endpoint.

Inject startup and restart faults

Delay the local certificate agent and present a database endpoint that accepts TCP but rejects authentication. Confirm the service does not report ready prematurely. Crash it once and measure recovery, then make it fail deterministically until the start limit is reached. Preserve the first error and exit status before resetting failure state. Rotate a credential file and show that every process uses the intended generation after the supported restart or reload path. Verify the process cannot write outside its managed state directory while it still performs its legitimate work.

Output
Synthetic acceptance
Readiness: withheld until credential and database checks pass
Crash loop: 4 starts within 12 minutes, then visible failed state
Credential: old and new generation recorded per process
Journal: first failure available after reboot within disk budget
Timer: 7 missed hourly windows processed from ledger
Socket: 47-request burst reconciled by request ID

Preserve evidence and scheduled work

Generate a controlled log burst, reboot, and query the previous boot's first settlement error by unit. Confirm the journal budget and central-forwarding behavior. Pause the host for seven hourly timer periods; when it returns, the job must process seven ledger windows even if the timer issues only a catch-up start. Crash after one window's effect but before its checkpoint and verify that the unique window ID prevents a second settlement.

Exercise the listener boundary

Run a separate socket-compatible worker. Start only its socket unit, send an initial request, then kill the service during 47 synthetic requests. Record accepted, timed-out, and retried connections, plus completed effect IDs. A socket reporting active while the service cannot handle a user path fails the check. Keep the final evidence as unit versions, commands, first-error logs, request ledger, and measured recovery intervals.

Common Mistakes

  • Do not infer readiness from unit activation alone.
  • Do not let a timer substitute for the durable work ledger.
  • Do not treat a listening socket as proof that a service can answer.

Connected lessons

devops
project
Storage details