A failed write can arise from mount options, a block-device error, a full filesystem, or filesystem corruption. The response must distinguish those causes before changing the mount. Kernel logs, the mount's current options, device health, and the application's first failed write form the initial evidence. Filesystem repair tools are type-specific and can modify metadata; running them against a live mounted filesystem can be unsafe or produce invalid findings. ext4 checking and XFS repair follow different procedures, particularly when an XFS log is dirty. Recovery should occur on a detached clone or during a controlled outage with a backup and a tested restore path.
Linux read-only filesystem incidents: preserve evidence before repair
Operational decision
A billing worker receives read-only errors on /srv/billing after a storage-path timeout. The responder fences writes, records the first kernel error and provider volume event, and verifies that its replica has the last committed invoice ID. On a disposable clone, they examine the filesystem type and replay or inspect metadata through the supported procedure, then compare recovered invoice IDs with the application ledger. A remount to read-write is rejected as a first response because it could mask an active device fault. When the real host returns, acceptance requires a successful application write and a second reader's confirmation; a writable mount flag alone is insufficient.
findmnt -no SOURCE,FSTYPE,OPTIONS /srv/billing
journalctl -k -n 80
lsblk -fCost and verification
Stopping writes costs availability but limits further damage while the cause is unknown. Repair can discard or alter metadata, so retain a pre-repair image and record exactly which utility and options were used. A successful filesystem check does not prove application-level consistency across files and database transactions. Measure lost writes, time to restore a healthy writer, and the verified application recovery point. If the underlying device continues returning I/O errors, a repaired filesystem on the same unhealthy path is not an acceptable return to service.
Common Mistakes
- Do not run a modifying filesystem repair against the mounted live volume.
- Do not use a force-log-clear option as a routine first step.
- Do not equate a read-write mount with recovered application data.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Volume snapshots: test application-consistent restore
- Write-gap ledgers: account for acknowledged work that never reached the standby
- Backups and disaster recovery: prove the restore path
- Linux disk pressure: explain missing space before deleting application data
- Linux filesystem expansion: follow the block device to the mounted filesystem
