GitOps compares live resources with declared intent. Restoring an old Deployment directly in the cluster may be overwritten by self-healing or the next sync; restoring a historical Application state can also conflict with automated sync. A rollback plan names the known-good rendered release, the immutable image digest, how intent is changed, which controller is paused if needed, and the condition for restoring ordinary reconciliation. Database schema and external side effects may not be reversible with the manifest.
GitOps rollback: return desired state and controller authority together
Operational decision
A reconciliation API deploys a bad config generation. The on-call operator records the currently applied source and artifact digests, then selects an earlier accepted render whose image remains pullable and whose data contract still matches the migrated database. If automation would immediately reapply the bad commit, the operator pauses or narrows sync with a time-limited incident record, updates desired state to the known-good revision, reviews the rendered deletion set, and resumes reconciliation. A synthetic reconciliation request must succeed before the incident is closed. Run a negative drill: patch the live Deployment to a previous image while leaving Git unchanged, then observe how the controller handles it. A direct patch is acceptable only as a bounded emergency action with the intended Git change queued and a known reversion path. Verify secret and ConfigMap generations, not just the container image; the earlier image may fail against the newer configuration. Preserve the failed release's evidence for investigation, and check that rollback artifacts and any admission statements still exist in the destination registry.
Reconciliation API rollback
Failed release: source, render, image, config generation
Candidate: accepted prior render and data compatibility
Control: sync pause owner and expiry if needed
Intent: revert or explicit prior revision in Git
Pre-sync: deletion and retained-resource review
Proof: controller observes target and user path passes
Exit: ordinary reconciliation restoredCost and verification
A rollback is at least one render, review, sync, rollout, and user probe; its elapsed time depends on image availability, scheduling, and data compatibility. Keeping prior images and evidence consumes registry storage, but deleting them can turn rollback into a rebuild under incident pressure. Measure emergency patches that lack a matching intent change, pause duration, rollback user-path time, and controller reapplication of the failed revision.
Common Mistakes
- Do not rely on a live kubectl patch while self-healing still targets the bad revision.
- Do not assume an earlier image is compatible with the current database schema.
- Do not forget to restore ordinary reconciliation after the incident.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- GitOps emergency changes: preserve one recorded desired state
- Rollback image retention: keep every approved fallback pullable
- Database change safety: expand, migrate, contract
- Release evidence: tie one deployed digest to one approval decision
