A supported Terraform backend may lock state during operations that could write it. The lock prevents two runs from racing to publish incompatible state snapshots. A failed runner can leave a lock that needs recovery, but an operator who clears a live writer's lock creates the same concurrency hazard the backend was meant to prevent. Lock handling depends on the backend; a visible lock error alone does not establish that the lock is stale.
Terraform state locks: distinguish a stale lease from an active writer
Operational decision
A production plan reports a lock ID after a CI runner disappears. Stop scheduled applies for that workspace, identify the lock owner and operation, check the runner and provider activity, and confirm whether a state write finished. Preserve the current state version and remote change record before taking action. Only the team that owns the failed operation should use the lock ID to release its own abandoned lock; the command below is an example of the final recovery step, not a first response. Replan after unlock and compare the new plan with the failed run's last known actions. If remote infrastructure changed but state did not, reconcile that object before any new apply. Alert on lock age, queue length, and repeated recovery requests, while keeping planned long operations from being mistaken for failures.
terraform force-unlock "6b8c1c29-4f71-42a0-bc53-79c70ab5d541"Cost and verification
A lock serializes writers and can delay deployments during a long apply. That wait is cheaper than repairing divergent state, which can involve manual remote inventory and data risk. Disabling locking to save minutes removes the single-writer guarantee. Force-unlock itself does not alter infrastructure, but the next writer can if the previous one was still running. Measure how long the stale lock blocked releases and whether a new plan reveals incomplete remote effects before declaring recovery complete.
Common Mistakes
- Do not run with locking disabled merely to bypass a busy workspace.
- Do not unlock another operator's active run.
- Do not resume an apply before checking whether the failed run changed remote objects.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Terraform state: shared ownership and safe plans
- Ambiguous cloud creates: reconcile before repeating a timed-out mutation
- Infrastructure drift: distinguish emergency repair from unauthorized change
- Git change control: small merges and protected branches
