Build the drill in an isolated cluster with synthetic receipt resources, two controller replicas, and a disposable storage driver. Capture API write latency, webhook call errors, Pod network assignment events, volume attach duration, watch relist count, lease holder, and completed receipt effects before changing any failure condition. Keep the state-store portion in a separate self-managed test environment; a managed service may not expose its quorum members.
Project: recover a degraded cluster control plane
Block writes and placement
Stop a test admission webhook and submit one safe workload through its exact match rule. Show whether the request is denied, times out, or enters without validation, then restore the webhook and audit the gap. Stop the current etcd leader in the isolated three-voter test and verify writes resume after election; never intentionally stop a second member from a shared cluster. Exhaust a disposable Pod address pool and a CSI attachment slot separately. Record the distinct Pod events and restore each resource before proceeding. Under node pressure, admit one high-priority payment Pod and prove that displaced batch work has a replay path.
Cluster control recovery gates
Admission: configured failure result and skipped-policy audit
State store: majority retained; API write verified after election
Networking: new Pod gets a usable address and serves a request
Storage: intended claim mounts without a second writer
CRD: old stored version removed only after migration
Watch: cache rebuilt after expired version
Lease: stale leader cannot commit a lower epochRestore controller continuity
Serve two versions of a synthetic InvoiceRun custom resource. Write one old-version object, switch the storage version, and show that a new-version read alone does not rewrite the stored object. Run a supported migration and verify storedVersions before removing the old served version. Disconnect the receipt controller until its resourceVersion expires; it must relist, compare the rebuilt object set, and converge without duplicate effects. Pause its leader so the standby acquires the Lease, then resume the old process. The effect store must reject its obsolete generation. Inject a lost provider response after a route write and confirm the current leader reconciles the same operation key.
Cost and verification
Record control-plane requests, extra compute, storage-driver calls, and recovery time. The acceptance record must include a successful synthetic customer request, one durable receipt effect per business key, no active old writer, and a current controller cache. Green Pod status, a Lease holder, or a new-version GET alone does not satisfy those checks. Clean up only the isolated test resources after retaining the evidence.
Common Mistakes
- Do not disable every admission rule to repair one webhook.
- Do not force-detach a volume from an uncertain writer.
- Do not assume a Lease change fences an external effect.
Connected lessons
- Admission webhook outage: choose a deliberate failure path
- etcd quorum: preserve a voting majority during maintenance
- Pod address capacity: diagnose network allocation before adding nodes
- CSI attach limits: verify storage placement as well as CPU placement
- CRD storage migration: retire an API version only after stored objects move
- Controller watches: recover from expired resource versions without a relist storm
- Pod priority: reserve recovery capacity without evicting the wrong work
- Leader leases: fence side effects after ownership changes
- DevOps projects
