Availability of the running service and freshness of desired-state input are different properties. A controller may retain its last successful source artifact and continue to compare or serve status while it cannot fetch the new revision. The incident response must record the last fetched revision, artifact age, source error, and whether any newer approved change is waiting. Do not interpret a stable cluster as proof that a security patch or rollback committed during the outage has been applied.
GitOps source outages: distinguish last-good operation from fresh deployment
Operational decision
An order-routing platform loses access to its chart registry during an incident. The current Pods remain healthy, so on-call staff avoid restarting them merely to clear a source error. They compare the controller's last successful artifact digest and source revision with the approved release queue and discover that a certificate fix is pending. The team checks whether the fix can be delivered through a preapproved alternate source path, including its identity and authorization, or whether the service can remain on the last-good version until the registry returns. In a disposable drill, revoke source credentials after one successful sync, submit a new desired revision, and prove the freshness alert fires even though workload health remains green. Then restore access and verify that the controller fetches and applies the intended revision once, without deploying an intermediate unreviewed head. If local cached chart bytes are used, record their digest; a cache hit is a deployable artifact only if its integrity and scope match the approved candidate. Define a maximum source staleness appropriate to the service's patch and recovery commitments, not one global threshold for every Application.
Order-routing source state
Last successful revision: recorded commit and artifact digest
Last fetch: timestamp and source identity
Pending approved revision: recorded separately
Running workload: health and image digest
Outage decision: hold last-good or use approved alternate
Recovery check: intended revision fetched and applied onceCost and verification
Source checks are cheap, but retaining local artifacts costs controller storage; aggressive eviction can remove the only reproducible last-good input during an upstream outage. Measure source artifact age, fetch failure duration, pending revision count, and time from restored connectivity to applied state. Alerts should distinguish 'serving last-good' from 'unable to deploy new intent' so operators do not create a second outage while fixing the first.
Common Mistakes
- Do not call a running workload current when its source has stopped updating.
- Do not restart healthy Pods solely because the source fetch failed.
- Do not deploy an unverified cache entry as a substitute for a pinned release.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- GitOps reconciliation: desired state and drift
- Registry replicas: prove the recovery region has the exact release image
- Certificate renewal: verify the served certificate after issuance
- Alert design: page on impact and include a first action
