Infrastructure drift is a difference between declared configuration, Terraform state, and the real provider object. It may result from an emergency console edit, another controller, provider defaults, or an unreviewed change. A Terraform plan can reveal differences, but it cannot tell operators whether the declared value or the live value is the correct one after an incident.
Infrastructure drift: distinguish emergency repair from unauthorized change
Operational decision
A storage team notices an encryption setting changed outside its reviewed configuration. Freeze applies to the affected state, preserve audit events, and determine who changed the setting and why. Run a refresh-backed plan in a read-only review context and inspect the exact proposed actions. The shell block records a plan for review but must not be used to apply without deciding ownership. If the live setting is a valid emergency repair, update code through review and confirm the next plan is clean. If the change is unauthorized, contain access first, then restore the approved value with a controlled apply. Check dependent services and data before declaring recovery; changing a property back may not undo data created during the drift interval.
terraform init -input=false
terraform plan -input=false -out=drift-review.tfplan
terraform show -no-color drift-review.tfplanCost and verification
Plan review consumes operator time and provider API calls, but an automatic revert can worsen an incident when live state carries a needed repair. A saved plan can include sensitive values; store and delete it under the same controls as state. Refresh may be delayed by provider consistency and does not detect every external effect. Monitor drift on a cadence that matches resource risk, and avoid running multiple controllers against the same setting without an explicit owner.
Common Mistakes
- Do not auto-apply a drift plan before understanding the live change.
- Do not leave a saved plan in an unrestricted CI artifact.
- Do not confuse a clean plan with proof that external data is intact.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Terraform state: shared ownership and safe plans
- Terraform address refactoring: move state without recreating infrastructure
- Incident response: contain impact, then learn
- Cloud cost and capacity: assign an owner to each recurring resource
Practice and check
Advanced follow-up
- Terraform import: adopt one existing object under one state owner
- Terraform ignore_changes: name the second owner of every ignored field
Continue with: Infrastructure prompts: quarantine state secrets and drift.
