A runbook automation turns a repeatable response procedure into a command or workflow. Automation removes operator typing errors but can also repeat a mistaken action across every target faster than a person would. Its safety comes from explicit preconditions, a limited target set, a dry-run or read-only preview when possible, and a postcondition that measures the user-facing result.
Runbook automation: put a stop gate before the irreversible step
Operational decision
A claims incident requires pausing a queue consumer to protect a failing database. Build a script that first reads the current cluster, namespace, Deployment UID, replica count, and queue age. Require an incident ID and a named target; reject a wildcard namespace. The shell block illustrates a read-only preflight and deliberately stops before mutation. In the complete runbook, show the expected new replica count and allow a bounded approval to scale only that Deployment. After the action, verify the targeted Pods stop fetching while existing work drains, and watch the upstream backlog. Set a maximum pause duration and a separate resume path. If a retry sees the deployment already paused, it should report that state instead of assuming its prior attempt failed. Exercise both success and partial-failure paths before using the script in production.
set -euo pipefail
incident_id=INC-47
namespace=claims
deployment=claims-queue-worker
test -n "$incident_id"
kubectl config current-context
kubectl get deployment "$deployment" -n "$namespace" -o jsonpath='{.metadata.uid} {.spec.replicas}'
kubectl get pods -n "$namespace" -l app=claims-queue-workerCost and verification
Preflight checks and bounded approvals add seconds during an incident, but prevent a typo from affecting a whole cluster. Dry-run support differs by operation and may not predict external side effects, so the script also needs a measured postcondition. Repeated polling and retries can overload an already failing API server; use limited attempts and timeouts. Keep the manual recovery procedure available if automation itself or its credentials fail.
Common Mistakes
- Do not make a read-only preflight appear to perform the mitigation.
- Do not accept an unbounded target selector or retry loop.
- Do not declare success until the intended user and queue effects are observed.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Incident response: contain impact, then learn
- Overload shedding: refuse excess work before latency collapses
- Queue consumers: acknowledgement, idempotency, and backlog
- Break-glass access: recover control without permanent privilege
