Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Terraform state locks: distinguish a stale lease from an active writer

Last updated: 1 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A supported Terraform backend may lock state during operations that could write it. The lock prevents two runs from racing to publish incompatible state snapshots. A failed runner can leave a lock that needs recovery, but an operator who clears a live writer's lock creates the same concurrency hazard the backend was meant to prevent. Lock handling depends on the backend; a visible lock error alone does not establish that the lock is stale.

Operational decision

A production plan reports a lock ID after a CI runner disappears. Stop scheduled applies for that workspace, identify the lock owner and operation, check the runner and provider activity, and confirm whether a state write finished. Preserve the current state version and remote change record before taking action. Only the team that owns the failed operation should use the lock ID to release its own abandoned lock; the command below is an example of the final recovery step, not a first response. Replan after unlock and compare the new plan with the failed run's last known actions. If remote infrastructure changed but state did not, reconcile that object before any new apply. Alert on lock age, queue length, and repeated recovery requests, while keeping planned long operations from being mistaken for failures.

bash
terraform force-unlock "6b8c1c29-4f71-42a0-bc53-79c70ab5d541"

Cost and verification

A lock serializes writers and can delay deployments during a long apply. That wait is cheaper than repairing divergent state, which can involve manual remote inventory and data risk. Disabling locking to save minutes removes the single-writer guarantee. Force-unlock itself does not alter infrastructure, but the next writer can if the previous one was still running. Measure how long the stale lock blocked releases and whether a new plan reveals incomplete remote effects before declaring recovery complete.

Common Mistakes

  • Do not run with locking disabled merely to bypass a busy workspace.
  • Do not unlock another operator's active run.
  • Do not resume an apply before checking whether the failed run changed remote objects.

Connected lessons

Practice and check

devops
operations
Storage details