Break-glass access is an emergency path to perform an operation when ordinary approval or identity systems cannot meet the incident deadline. It is a deliberately exceptional privilege, not a second everyday administrator account. The path needs a scope, independent audit record, expiration, and a way to disable it once normal control returns. A recovery design that relies on the failed identity provider for every step may not work during an identity outage.
Break-glass access: recover control without permanent privilege
Operational decision
A cluster incident blocks the usual deployment identity while a faulty admission policy prevents replacement Pods. An incident commander authorizes a named operator to use an emergency role for the affected namespace. The text block specifies the evidence to retain; the exact credential issuance and access checks depend on the platform. Before use, test that the role can patch the single policy binding but cannot read unrelated customer secrets. During use, log the actor, incident, command, object diff, and result to a separately protected audit destination. Revoke the credential on recovery, compare live state with Git, and file the minimal corrective change. Drill the path periodically with a harmless resource in a disposable environment. Avoid shared static passwords whose use cannot be attributed.
Emergency cluster access gate
Incident: named response record
Actor: individually identified operator
Scope: one namespace policy binding
Duration: short automatic expiry
Audit: independent command and object-diff record
Exit: revoke credential, reconcile desired state, review useCost and verification
An emergency role adds a sensitive credential and a monitoring burden, but can shorten recovery when regular control fails. Too little scope can make it useless; too much makes compromise expensive. Pre-test its exact recovery action and deny unrelated reads. Automatic expiry and post-incident reconciliation reduce persistent privilege, but they do not undo changes made while the role was active. Keep a second operator or reviewer available for high-impact changes when the incident permits it.
Common Mistakes
- Do not make the only recovery path depend on the broken identity service.
- Do not share an unattributed emergency credential.
- Do not leave emergency privilege active after normal access returns.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes RBAC: bind one service account to one job
- GitOps emergency changes: preserve one recorded desired state
- Credential incident response: revoke access before rebuilding trust
- Incident response: contain impact, then learn
Practice and check
Cloud authority follow-up
- Cloud policy decisions: trace every authorization layer
- Audit log integrity: verify delivery and digest continuity
