This drill exercises a synthetic claims upload path across a host, a public name, and an asynchronous recovery copy. Use a disposable cluster, test records, and a dedicated object bucket. Capture the baseline request latency, cgroup CPU counters, open descriptors, inode availability, clock offset, queue age, and object-copy status before making changes.
Project: recover a host and protocol failure chain
Induce and distinguish failures
Apply a low CPU limit to one test Pod and show throttling grows while node capacity remains available. Open and close enough test connections to reveal descriptor growth, then create small disposable spool files until an inode alert fires without deleting unprocessed work. Offset the clock in an isolated host or simulated time environment and observe a test token validity failure; restore synchronization before continuing. Query a hostname before its record exists, create it, and compare authoritative and recursive answers. During a rollout, hold a long upload on a reused connection and verify its result after the old endpoint stops receiving new traffic.
Claims host-protocol drill
CPU: throttled time delta and p95 latency
Files: descriptor count and inode use before/after cleanup
Time: offset and test-token acceptance
DNS: cached missing-name answer and eventual success
Drain: one long upload completes exactly once
Queue: lease expiry yields one durable effect
Recovery: required object version copied and parsedRecover delayed work
Hold one attachment transcode beyond its initial visibility window. Renew while progress continues, then crash after durable output but before acknowledgement; the next delivery must reuse the original operation key. Write an archive object and inspect the exact version in both source and recovery stores. Delay replication and confirm the disaster-recovery gate refuses a cutover while the required version is pending. Once copied, compare digest and application parse result before marking the recovery point complete.
Cost and verification
The drill spends test compute, queue deliveries, object-copy traffic, and diagnostic time. Record the added latency, descriptor and inode recovery, DNS negative-cache duration, interrupted connection count, duplicate effect count, and object-copy lag. A green Pod status or enabled replication rule alone does not satisfy the acceptance criteria. Remove only disposable resources after the evidence is stored under the team's retention policy.
Common Mistakes
- Do not treat free disk bytes as free inodes.
- Do not assume an endpoint change closes existing connections.
- Do not acknowledge a queue message before its output is durable.
Connected lessons
- CPU throttling: distinguish a quota ceiling from node contention
- File descriptor exhaustion: find the leak before raising the limit
- Inode exhaustion: diagnose a full filesystem with free bytes
- Clock skew: verify time before debugging credentials and leases
- DNS negative caching: avoid creating a name after clients already asked for it
- Connection draining: let in-flight work finish while new traffic moves
- Queue visibility leases: prevent overlapping workers on one message
- Object replication: verify the exact recovery object arrived
- DevOps projects
