Use an isolated receipt API, a ledger stub that can delay or commit requests, two application replicas, a test load balancer, and an edge cache with separate public and account-specific routes. Issue synthetic requests with stable operation IDs. Before faults, record the deadline budget, active HTTP/2 streams, health-check target state, client peer addresses, payout count, mirror destinations, and effective response cache headers. Do not connect the shadow path to real payment, email, or account services.
Project: release a service through connection and cache failures
Interrupt traffic safely
Delay the ledger beyond a 430-millisecond receipt lookup budget and prove cancellation reaches the downstream call. Start concurrent HTTP/2 requests, mark one replica unready, and ask it to shut down gracefully; measure completed accepted streams, rejected new streams, and the hard termination boundary. Break the receipt handler while leaving the socket open and compare load-balancer target health with a synthetic user request. Change a gateway address, then record when the resolver, client connection pool, and last old-version response each converge. Pace pool rotation to avoid a synchronized reconnect surge.
Service traffic acceptance gates
Deadline: downstream call stops after parent budget expires
Drain: accepted streams complete or fail within recorded grace window
Health: application fault withdraws target; synthetic path reports impact
Cutover: no unexplained old-version responses after pool rotation
Retry: one payout per operation ID after lost response
Mirror: no real side effect or private payload escape
Assets: both current and rollback digest objects remain fetchable
Cache: stale public catalog bounded; private payout never sharedReplay and cache with evidence
Commit a synthetic payout, discard its client response, and retry with the same operation key; verify one ledger effect and the same durable result. Reuse that key with a changed amount and require rejection. Mirror a fraction of receipt reads to a test-only backend, then inspect its storage and outbound connections for side effects. Publish a digest-named bundle before switching the document reference, keep the old bundle through rollback, and verify both from separate edge locations. Stop the public catalog origin and test bounded stale delivery; repeat with two private test accounts to prove payout responses never cross identities or appear in a shared cache.
Cost and verification
Record extra overlap replicas, handshake rate, shadow egress, idempotency record growth, old-bundle storage, and origin load from uncached private requests. The drill passes only when a lost response does not duplicate a payout, the old endpoint stops serving after the agreed window, and the effective edge policy matches the route's data sensitivity. A green listener, a successful DNS lookup, and an unchanged cache-hit ratio are not substitutes for these checks.
Common Mistakes
- Do not equate target readiness with drained HTTP/2 streams.
- Do not retry a committed payment without a durable operation key.
- Do not mirror writes into real downstream systems.
Connected lessons
- gRPC deadlines: spend one request budget across every downstream call
- HTTP/2 drain: let existing streams finish while new calls move away
- Load-balancer health: a green socket is not a working transaction
- Client connection pools: retire old endpoints after a DNS or rollout change
- Idempotency keys: reconcile an accepted write before repeating it
- Shadow traffic: test a new backend without duplicating live side effects
- Versioned edge assets: make long cache lifetimes safe through content identity
- Edge cache fallback: serve stale public data without leaking private responses
- DevOps projects
