Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Linux I/O latency: distinguish device wait from an application queue

Last updated: 2 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

The request may wait in an application pool, filesystem, block queue, controller, network-attached volume, or database lock before a device reports completion. iostat extended fields describe observed device requests and queueing over a sampling interval; await includes time spent waiting for the requests represented there, and it is not the same as an application's full request latency. High utilization alone is ambiguous on concurrent storage devices. Compare read and write behavior, throughput, queue depth, device errors, and application tail latency during the same interval, then test a change one bottleneck at a time.

Operational decision

A media transcoder's metadata writes slow from 18 ms to 430 ms at peak load. The operator samples iostat over successive intervals instead of trusting its first since-boot average, correlates the affected device with the actual metadata mount, and checks application queue wait separately. The volume shows rising write await and queue length, while reads remain steady. A test cohort reduces batch flush concurrency from 23 to 9, after which device queue depth and the 99th-percentile metadata write time fall together. If the application remains slow while device waits return to baseline, the investigation moves to locks or thread-pool saturation instead of purchasing larger disks without evidence.

bash
findmnt -no SOURCE /srv/media-meta
iostat -xz 2 4
journalctl -k -n 60

Cost and verification

Longer sampling windows smooth spikes and may conceal a short stall; very short windows can produce noisy rates. Tracing every I/O on a busy host has CPU and log costs, so reserve it for a bounded investigation after inexpensive counters identify the device and interval. A smaller batch may reduce throughput even as tail latency improves; assess both against the service objective. For network-backed volumes, provider throttling and path failures can raise latency without high local device utilization. Preserve the before-and-after workload rate when judging an optimization.

Common Mistakes

  • Do not interpret the first iostat line as a fresh interval sample.
  • Do not diagnose from utilization percentage without queue and latency data.
  • Do not compare device and application latency from different time windows.

Connected lessons

Practice and check

devops
linux
storage-operations
Storage details