Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: run a service resilience game day

Last updated: 5 Oct 20269 min read
project
AdvancedBy AITrove Editorial

A resilience game day is a controlled exercise that tests whether a service meets a stated user and recovery objective during planned faults. Use a disposable two-zone environment, synthetic claims, and credentials created only for the exercise. State the expected transaction success rate, maximum request latency, and incident response clock before changing the system.

Prepare the system

Run four claims API replicas with a PDB that needs three ready, spread them across zones, and implement bounded graceful shutdown. Give the API a stable idempotency key for a claim submission. Record the image digest, the current deployment revision, the synthetic claim IDs, and how to reset the environment. Install a log-loss signal and a client-side user-path probe. Do not use a real production token as the credential-leak prop.

Output
Claims game-day record
Scenario A: drain one node while a slow request is in flight
Scenario B: cut signer response after commit; retry same claim key
Scenario C: expose disposable CI token in private test log
Evidence: PDB status, client result, claim count, token denial
Stop gate: user error threshold or unexpected data write
Owner: named operator and observer

Run and recover

Drain one node through the Eviction API and watch both PDB status and client errors. Interrupt the remote signer after a durable write, retry with the same idempotency key, and verify that only one claim was created. Place the disposable token in a restricted test log, then revoke it at the issuer, test old-token denial, create a scoped replacement, and resume one test deployment. Stop immediately if the exercise creates unexpected writes or loses visibility. Capture actual recovery times and every manual decision, including a stalled drain.

Cost and verification

The exercise consumes temporary cluster capacity, log storage, and operator time. That cost buys direct evidence about the system's failure path, but only for the conditions actually tested. Record whether the PDB allowed eviction, whether the client received one durable result, whether the old token failed, and which observation was missing. Remove disposable resources and credentials after preserving the approved evidence.

Common Mistakes

  • Do not call a successful drain a user-path test.
  • Do not infer idempotency from a single request without an uncertain first result.
  • Do not remove a leaked test log in place of revoking its token.

Connected lessons

devops
project
Storage details