This drill tests a synthetic payout service at points where a green Pod or successful HTTP response can hide lost work. Use a disposable database, event log, collector, and cluster. Record the initial primary and standby positions, consumer offsets, active time series, trace drop rate, and current access policy before injecting failure.
Project: drill replica lag, replay, and incident control
Induce bounded faults
Delay standby replay, acknowledge a payout write, and read from the standby. Record whether the old value appears and route the confirmation read to the primary. Kill a consumer after its ledger effect commits but before its offset commits; verify a replacement retries the same record without a second ledger effect. Add many distinct synthetic account IDs to request traffic and check that metric-series count remains bounded. Generate one payment timeout and confirm the trace selection path retains its spans without sensitive data.
Failure-boundary evidence
Primary write ID and standby replay position
Original event ID, partition, offset, one ledger effect
Series count before and after unique account IDs
Injected timeout trace ID and collector drop count
Emergency actor, target, expiry, and observed postcondition
Recovery: user transaction succeeds after mitigationRecover control
Create a disposable resource whose cleanup controller is temporarily unavailable, request deletion, and inspect its finalizer. Restore the controller and verify cleanup finishes without manually stripping the field. Rehearse a read-only incident preflight for pausing the payout consumer, then perform the mutation only through a reviewed, target-limited procedure. Restore normal operation and revoke any temporary role. Record a policy exception only if the baseline admission rule would otherwise block the recovery, and prove that the exception no longer exists at its expiry.
Cost and verification
The exercise consumes extra database, broker, collector, and operator capacity. Measure replay lag, duplicate-effect count, metric-series growth, dropped spans, and deletion completion time. A successful script exit is insufficient if the wrong Deployment was changed or a customer-facing payout remains stale. Retain the incident evidence under the team policy and remove synthetic records after verification.
Common Mistakes
- Do not confuse the committed consumer offset with a durable payout effect.
- Do not remove a finalizer before completing its cleanup.
- Do not call a sampled trace a complete account of all traffic.
Connected lessons
- Replica lag: define when a read is allowed to be stale
- Consumer rebalances: preserve ordering and effect ownership
- Metric cardinality: keep observability usable during a surge
- Trace sampling: retain useful failures without flooding storage
- Stuck finalizers: finish cleanup before removing the guard
- Runbook automation: put a stop gate before the irreversible step
- Policy exceptions: make a temporary bypass expire and prove its scope
- Break-glass access: recover control without permanent privilege
- DevOps projects
