A green health endpoint can hide an empty image registry, an expired certificate, a missing signing key, a quota ceiling, or a database replica too far behind the agreed data-loss bound. Admission is a pre-cutover decision: test the dependency graph and identify which checks are hard gates, which allow degraded service, and who may accept a bounded exception. Record actual values at decision time because a readiness result from yesterday does not describe the current outage.
Standby admission: prove the recovery region can accept real work
Operational decision
A shipment-tracking platform keeps a reduced-size standby. Before any traffic move, verify the deployment digest is pullable from a clean node, the alternate-region secrets and certificates decrypt, the database is on the expected timeline, queue consumers are paused or scoped, and image, compute, database, IP, and load-balancer quotas can cover a measured intake rate. Compare the projected rate with requests and autoscaling limits; a warm standby that serves one test request may still collapse under the first peak. Test outbound dependencies such as tax calculation and email by making a scoped synthetic request, while preventing the probe from sending real messages. A key needed to decrypt retained backups receives a separate gate from the key used by current traffic. If one nonessential analytics pipeline is absent, publish the resulting service mode and duration; if order writes cannot be fenced, block promotion. The admission sheet records owner, evidence time, observed value, threshold, decision, and a link to the repair action. Re-run the checks after scaling because new nodes may require a registry path and credentials the original warm nodes already cached.
Shipment standby admission
Artifact: tested digest pulls on clean node
Data: replay point inside accepted loss window
Writes: old primary fenced; new epoch available
Identity: service tokens and certificates valid
Keys: current data and retained restore points decrypt
Capacity: quota plus measured intake headroom
Queues: producer and consumer ownership defined
Decision: hard gate, degraded mode, or blockCost and verification
A full rehearsal costs a clean-node pull, scoped dependency probes, and capacity tests; those checks consume traffic and provider quota, so run destructive or costly probes in a disposable environment. With D dependencies, the admission inventory is O(D), but the slowest critical dependency may dominate time to readiness. Track failed gates, stale evidence age, first-hour saturation, and clean-node pull success instead of counting only green health endpoints.
Common Mistakes
- Do not accept a standby from a single HTTP health response.
- Do not count cached images or credentials as proof a new node can start.
- Do not promote when the old writer may still accept transactions.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Provider quota preflight: reserve capacity for rollback and recovery
- Node autoscaling: make pending Pods schedulable before traffic rises
- Registry replicas: prove the recovery region has the exact release image
- Encryption key rotation: keep old data decryptable during recovery
