A GitOps controller repeatedly compares declared configuration with live resources and may restore drift. A manual cluster edit can be overwritten if self-healing is enabled. Pausing reconciliation can keep an urgent fix in place, but it also suspends correction of unrelated drift. The incident procedure needs a bounded pause, an owner, a repository change, and a verified return to reconciliation.
GitOps emergency changes: preserve one recorded desired state
Operational decision
A payments API needs an immediate route correction after an unsafe policy reaches production. First identify the application and exact object the controller owns; record its current revision and health. If a manual correction is necessary, pause or narrowly disable synchronization under an incident record, apply the smallest safe change, and measure the user path. Create the matching repository change before resuming automated sync. The checklist below is a decision record, not a command sequence for a particular controller. Compare live and desired objects after reconciliation restarts; the incident is not closed because kubectl accepted a patch. Revert the temporary pause and verify that a later safe change still deploys. If the incident involves a compromised repository, do not treat that repository as trusted desired state until access is contained.
Emergency GitOps change record
Application: payments-api
Live revision and affected object: captured before edit
Sync pause owner and expiry: named incident commander
Manual change: narrow object patch with user-path check
Repository repair: reviewed matching revision
Exit gate: sync resumed, live equals desired, service healthyCost and verification
A pause reduces automated interference but raises drift risk for every object within its scope. Keep it short and visible. A direct patch can be faster than a repository review, yet a forgotten patch will either disappear or remain untracked. Test the recovery path under a disposable application before production, including what happens when a sync window or policy blocks the intended change. Retain the incident record with the exact revisions and operator actions.
Common Mistakes
- Do not leave synchronization paused after the incident.
- Do not rely on a manual patch surviving self-healing.
- Do not merge an unreviewed repository change only to match a live mistake.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- GitOps reconciliation: desired state and drift
- Incident response: contain impact, then learn
- Kubernetes admission policy: reject an unsafe workload before scheduling
- Progressive delivery: canary checks and rollback
