A Kubernetes watch can end, and an old resourceVersion may no longer be available. A controller that insists on continuing from that version can miss changes or loop on a Gone response. A safe controller rebuilds its local view from a list, resumes watching from the returned version, and reconciles desired effects idempotently. Large, simultaneous relists can strain the API server and its backing store.
Controller watches: recover from expired resource versions without a relist storm
Operational decision
A custom receipt controller disconnects during an API-server restart. Keep its last observed resourceVersion with diagnostic context, then treat a version-expired response as a signal to discard the stale cache and relist the relevant resource scope. Use paginated reads where supported, respect API retry guidance, and add jitter so many replicas do not relist at the same instant. The policy fragment describes the loop; it is not a complete client library. After the list, reconcile every observed object and watch from the list's returned version. If a delete happened during the outage, compare the rebuilt set with the prior cache and run finalization only under the controller's idempotency guard. Test repeated disconnects and a restart while a resource changes. Watch count, relist duration, API rejection rate, and reconciliation lag reveal whether recovery is healthy. A connected watch alone does not prove the cache is complete.
Receipt controller watch contract
LIST scoped resources; retain returned resourceVersion
WATCH from that version; reconcile each event idempotently
On expired version: discard stale cache, jitter, relist, compare sets
On API throttle: honor retry guidance and total deadline
Acceptance: current object set and durable effects match desired stateCost and verification
A relist consumes API and memory capacity proportional to the scoped object set. Narrow selectors help only when they still cover every object the controller owns. Unbounded watch retries can starve the API server during recovery; a fixed synchronized delay can create another burst. Reconciliation must tolerate events repeated after relist. Measure progress from last completed reconciliation, not simply the age of the network connection.
Common Mistakes
- Do not retry an expired resourceVersion forever.
- Do not accept a newly connected watch as proof the cache is complete.
- Do not let every controller replica relist simultaneously without a rate budget.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- GitOps reconciliation: desired state and drift
- Cloud API throttling: keep infrastructure changes inside a request budget
- Stuck finalizers: finish cleanup before removing the guard
- Observability: join metrics, logs, and traces
