Skip to content
AITroveRead. Build. Understand.
Make this comfortable

etcd quorum: preserve a voting majority during maintenance

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A replicated etcd cluster needs a voting majority to commit changes. In a three-member cluster, one unavailable member can be tolerated; taking a second offline loses write availability. A leader election can briefly pause writes even when a majority survives. Existing application Pods may keep serving traffic while Kubernetes cannot persist a new desired state, making a healthy data plane misleading during control-plane failure.

Operational decision

Before planned maintenance on a self-managed cluster, inventory voting members, their zones, current leader, endpoint health, and a restorable snapshot. The sample uses an already configured, authenticated etcdctl client; it reads status and does not alter membership. Drain or replace only one member at a time, wait for the replacement to catch up, and verify API write health before continuing. If a majority is lost, stop automation that keeps issuing writes. Recover the missing members when possible; if permanent loss forces snapshot recovery, isolate the old cluster so it cannot later rejoin as a competing state source. Reconcile the snapshot's recovery point with external systems before resuming controllers. Managed-cluster users may not have direct etcd access; use the provider's control-plane health evidence and escalation path instead. A snapshot file existing is insufficient: practice restore into an isolated environment and inspect representative cluster objects.

bash
etcdctl --endpoints="$ETCD_ENDPOINTS" endpoint status --cluster --write-out=table
etcdctl --endpoints="$ETCD_ENDPOINTS" endpoint health --cluster

Cost and verification

Extra members increase infrastructure cost, and adding a member to a degraded cluster can change the quorum requirement before it is healthy. Cross-zone placement lowers correlated failure risk but increases network dependence and write latency. Repeated full snapshots consume storage and I/O. Monitor leader changes, failed proposals, disk commit latency, and Kubernetes API write results together. Do not infer quorum from process liveness or a single reachable endpoint.

Common Mistakes

  • Do not restart two of three voting members in one maintenance wave.
  • Do not add a member to a broken cluster without calculating the new quorum requirement.
  • Do not reconnect an old cluster after snapshot recovery without explicit isolation and reconciliation.

Connected lessons

Practice and check

devops
operations
Storage details