DevOps connects the code change, the deployed artifact, the observed user result, and the recovery path. Work through one service and retain the commit, digest, deployment, and incident evidence that ties those stages together.
Delivery foundations
Follow a change through review, verification, and promotion.
- DevOps: map a change from commit to recovery
- DevOps feedback loops, ownership, and release boundaries
- Git change control: small merges and protected branches
- Continuous integration: test the merge candidate
- Immutable artifacts and release provenance
Build and secure
Package one verified artifact and narrow job authority.
- Container builds: small runtime, explicit privilege
- GitHub Actions: narrow tokens and cloud trust
- Secrets and configuration across the delivery path
Infrastructure and release
Control desired state, rollout capacity, compatibility, and rollback.
- Terraform state: shared ownership and safe plans
- GitOps reconciliation: desired state and drift
- Kubernetes Deployment: rolling update capacity
- Kubernetes probes: startup, readiness, and liveness
- Progressive delivery: canary checks and rollback
- Database change safety: expand, migrate, contract
Operate and recover
Measure user impact, respond to incidents, and prove restoration.
- Observability: join metrics, logs, and traces
- SLOs and error budgets: turn reliability into a decision
- Incident response: contain impact, then learn
- Backups and disaster recovery: prove the restore path
Apply and check
Hosts and edge
Locate the failing process, name, certificate, or proxy boundary.
- Linux service diagnostics: process, socket, and journal
- DNS, TLS, and reverse-proxy failure boundaries
- Configuration management: converge hosts without surprise restarts
Build trust
Isolate runners and verify what a release contains and where it was built.
- CI runner isolation: treat repository code as untrusted
- Software supply chain: SBOM and provenance at admission
Cluster capacity and routing
Size Pods, scale from demand, and trace the route to a ready backend.
- Kubernetes requests and limits: schedule for real load
- Horizontal autoscaling: choose a signal tied to demand
- Kubernetes Service discovery: selectors, endpoints, and DNS
- Persistent storage: PVC lifecycle and data ownership
Cluster access and operations
Limit traffic and API authority; page on impact and process queued work safely.
- Kubernetes NetworkPolicy: permit only required flows
- Kubernetes RBAC: bind one service account to one job
- Alert design: page on impact and include a first action
- Queue consumers: acknowledgement, idempotency, and backlog
- Capacity and load tests: identify the next bottleneck
Release packaging and routes
Review the rendered release, route attachment, and runtime exposure.
- Helm release review: render before applying
- Gateway API routing: accepted route versus working request
- Feature flags: stop exposure without pretending code vanished
Trusted build inputs
Keep caches disposable and verify admission evidence before deployment.
- CI dependency caches: speed without hidden build inputs
- Kubernetes admission policy: reject an unsafe workload before scheduling
Cross-service evidence
Connect traces to user impact without propagating private data.
Infrastructure ownership and recovery
Keep module, cost, and failover contracts explicit.
- Terraform modules: small interfaces and explicit state owners
- Cloud cost and capacity: assign an owner to each recurring resource
- Multi-region failover: define write ownership before moving traffic
Availability during maintenance
Preserve service capacity while nodes and Pods leave the cluster.
- Pod disruption budgets: make node drains measurable
- Topology spread: keep replicas out of one failure domain
- Graceful Pod shutdown: stop accepting work before exit
- Cluster upgrade drill: preserve a path through each version step
Bounded background and remote work
Give schedules, retries, and contract changes explicit failure behavior.
- CronJob schedule safety: missed runs, overlap, and replay
- Retries and timeouts: bound the cost of a failed request
- API compatibility windows: release consumers and producers safely
Runtime evidence and response
Restrict containers, retain useful signals, and close credential exposure.
- Container runtime restrictions: remove privileges a service does not use
- Log pipelines: preserve incident evidence without ingesting secrets
- Credential incident response: revoke access before rebuilding trust
Identity and deployment ownership
Scope workload credentials and record one approved release.
- Service account tokens: mount only when the workload needs Kubernetes API access
- Federated workload identity: replace standing cloud keys with scoped trust
- GitOps emergency changes: preserve one recorded desired state
- Release evidence: tie one deployed digest to one approval decision
Infrastructure state and recovery
Refactor state deliberately and test recoverable storage.
- Terraform address refactoring: move state without recreating infrastructure
- Volume snapshots: test application-consistent restore
- Infrastructure drift: distinguish emergency repair from unauthorized change
Edge and capacity boundaries
Prove certificate rotation, egress control, and overload behavior.
- Certificate renewal: verify the served certificate after issuance
- Egress policy and DNS: restrict destinations without breaking name resolution
- Overload shedding: refuse excess work before latency collapses
Stateful change and replay
Stage persistent workloads and make delayed work recoverable.
- StatefulSet rollout safety: preserve identity and data while updating
- Database backfills: checkpoint progress without racing live writes
- Dead-letter replay: recover failed messages without repeating their effects
- Secret rotation rollout: update the issuer, consumer, and active connections
Capacity and release signals
Bound namespace demand and use user-facing evidence to advance a release.
- Namespace quotas: reserve room for a safe rollout
- SLO burn-rate alerts: page on budget consumption, not isolated spikes
- Canary analysis: compare a small cohort without hiding its failures
Follow-through and review environments
Close incident actions and temporary deployment access with evidence.
- Incident reviews: turn a timeline into tested corrective work
- Ephemeral environments: keep preview access and cost bounded
- Dependency patch campaigns: update, test, and prove the running image
Trust and shared-resource limits
Authenticate peers and budget database sessions across replicas.
- Mesh identity: require encrypted peers and narrow service access
- Database pool pressure: bound waiting before the database collapses
Node and event boundaries
Trace involuntary evictions, delayed capacity, and retained message contracts.
- Node pressure eviction: trace lost Pods to exhausted local resources
- Node autoscaling: make pending Pods schedulable before traffic rises
- Event schema evolution: release consumers before new event shapes
External release evidence
Control DNS overlap, deployment order, and user-path probes.
- DNS cutovers: budget for resolver caches and mixed destinations
- CI deployment concurrency: serialize mutations without losing change evidence
- Synthetic transactions: measure the route a user actually takes
Data and replay failure modes
Protect acknowledged writes and partition ownership through failure.
- Replica lag: define when a read is allowed to be stale
- Consumer rebalances: preserve ordering and effect ownership
Telemetry under pressure
Control series and trace volume without losing failure evidence.
- Metric cardinality: keep observability usable during a surge
- Trace sampling: retain useful failures without flooding storage
Control-plane recovery
Complete deletion and constrain exceptional incident authority.
- Stuck finalizers: finish cleanup before removing the guard
- Policy exceptions: make a temporary bypass expire and prove its scope
- Break-glass access: recover control without permanent privilege
- Runbook automation: put a stop gate before the irreversible step
Recovery authority and retained data
Prevent competing writers and preserve decryptable recovery points.
- Failover fencing: prevent two writable database leaders
- Immutable backup retention: protect recovery copies from deletion
- Encryption key rotation: keep old data decryptable during recovery
Ordered deployment and capacity
Stage API changes, replica behavior, and cache-origin work.
- GitOps rollout order: install API definitions before dependent objects
- HPA stabilization: prevent replica oscillation without masking demand
- Cache stampedes: bound origin work when popular keys expire
Edge fairness and image variants
Bound abusive traffic and verify every platform in a promoted image.
- Edge rate limits: reject abuse without penalizing shared networks
- Multi-architecture images: verify every platform behind one tag
Host exhaustion and time
Diagnose quota, descriptor, inode, and clock failures before changing the service.
- CPU throttling: distinguish a quota ceiling from node contention
- File descriptor exhaustion: find the leak before raising the limit
- Inode exhaustion: diagnose a full filesystem with free bytes
- Clock skew: verify time before debugging credentials and leases
Name and connection lifetime
Account for cached absence and persistent connections during rollout.
- DNS negative caching: avoid creating a name after clients already asked for it
- Connection draining: let in-flight work finish while new traffic moves
Delayed work and recovery objects
Guard long-running queue work and verify copied recovery data.
- Queue visibility leases: prevent overlapping workers on one message
- Object replication: verify the exact recovery object arrived
Cloud egress and control-plane limits
Budget outbound connections, API requests, and deployment surge capacity.
- NAT port pressure: find the shared outbound ceiling
- Cloud API throttling: keep infrastructure changes inside a request budget
- Provider quota preflight: reserve capacity for rollback and recovery
Backlog and recovery traffic
Use completion age and bounded retry traffic to protect recovery.
- Queue-age scaling: target completion time, not only queue length
- Retry amplification: assign one owner for each failed operation
- Circuit breaker recovery: probe capacity without reopening a flood
Isolation and ambiguous mutations
Separate dependency capacity and reconcile uncertain creates.
- Dependency bulkheads: stop one slow path consuming every worker
- Ambiguous cloud creates: reconcile before repeating a timed-out mutation
API write and state-store availability
Test admission failure, quorum, and version migration as control-plane dependencies.
- Admission webhook outage: choose a deliberate failure path
- etcd quorum: preserve a voting majority during maintenance
- CRD storage migration: retire an API version only after stored objects move
Scheduling limits beyond compute
Account for Pod addresses, volume slots, and the cost of preemption.
- Pod address capacity: diagnose network allocation before adding nodes
- CSI attach limits: verify storage placement as well as CPU placement
- Pod priority: reserve recovery capacity without evicting the wrong work
Controller continuity
Rebuild watches safely and fence effects after leader changes.
- Controller watches: recover from expired resource versions without a relist storm
- Leader leases: fence side effects after ownership changes
Build identity and private inputs
Separate untrusted artifacts, deployment claims, build secrets, and package sources.
- Untrusted CI artifacts: separate a pull-request test from promotion
- CI OIDC claims: bind cloud access to the exact deployment job
- Build secrets: keep credentials out of layers and exported caches
- Private package names: prevent a public registry winning resolution
Registry identity and signing trust
Bind a release to exact bytes and rotate trust without skipping verification.
- Registry tag mutation: reject release decisions based on a moving pointer
- Signature trust rollover: rotate verification roots without disabling admission
Recovery image availability
Prove exact digests remain pullable in every required region and rollback window.
- Registry replicas: prove the recovery region has the exact release image
- Rollback image retention: keep every approved fallback pullable
Change streams and durable event intent
Budget retained WAL, hand a snapshot into streaming, and publish committed state safely.
- Replication slots: bound retained WAL before a consumer outage fills the disk
- CDC snapshot handoff: prove that initial rows and later changes form one history
- Transactional outbox: commit business state and event intent together
Online schema and storage maintenance
Control index work, DDL waits, vacuum pressure, and partition retirement.
- Concurrent index builds: verify validity after the command exits
- DDL lock budget: fail a migration before it stalls production traffic
- Vacuum horizon: find transactions holding old row versions alive
- Partition retention: detach, prove the archive, then delete
Restore timeline acceptance
Prove the selected recovery point against business effects.
Truth in service measurements
Detect missing targets and count requests at the latency boundary.
- Scrape staleness: separate a failed target from a missing target
- Latency histograms: put a bucket at the actual SLO threshold
Telemetry transport under failure
Budget remote-write and Collector backlogs before data is lost.
- Remote-write backlog: budget the gap between local samples and long-term storage
- Collector export queues: measure the failure budget before telemetry is dropped
Trace and alert decisions
Keep traces together and route root-cause pages to their owner.
- Tail sampling at scale: keep every span of a trace with one decision maker
- Alert routing and inhibition: suppress symptoms without silencing the cause
Monitoring trust boundaries
Test redundant monitoring and remove sensitive fields before export.
- Monitoring redundancy: keep duplicate collectors without double-counting requests
- Telemetry redaction: remove sensitive fields before an exporter or sampler sees them
Request and connection lifetime
Carry one deadline across calls and drain multiplexed connections.
- gRPC deadlines: spend one request budget across every downstream call
- HTTP/2 drain: let existing streams finish while new calls move away
Routing truth during change
Compare health decisions with user requests and retire stale client pools.
- Load-balancer health: a green socket is not a working transaction
- Client connection pools: retire old endpoints after a DNS or rollout change
Safe traffic duplication
Reconcile ambiguous writes and contain mirrored requests.
- Idempotency keys: reconcile an accepted write before repeating it
- Shadow traffic: test a new backend without duplicating live side effects
Edge freshness and privacy
Pin immutable asset bytes and bound stale public responses.
- Versioned edge assets: make long cache lifetimes safe through content identity
- Edge cache fallback: serve stale public data without leaking private responses
Dependency and adoption integrity
Pin executable dependencies and adopt existing objects under one owner.
- Terraform provider locks: review the executable dependency
- Terraform import: adopt one existing object under one state owner
Replacement and lock recovery
Check coexistence capacity and recover state only after the writer stops.
- Terraform replacement: prove old and new can coexist
- Terraform state locks: distinguish a stale lease from an active writer
State secrecy and instance identity
Control sensitive artifacts and keep collection keys stable.
- Terraform state secrets: redaction is not removal
- Terraform collection keys: keep resource identity stable through reorderings
Shared control and handoff
Name the owner of ignored fields and remove state addresses deliberately.
- Terraform ignore_changes: name the second owner of every ignored field
- Terraform state removal: hand off an object without deleting it
Packet and flow capacity
Find size-dependent failure and exhausted node flow tables.
- Path MTU: find the packet size that breaks an otherwise healthy route
- Connection tracking: separate host flow exhaustion from application limits
Connection lifetime and identity
Align idle timers and trust client addresses only across known proxies.
- Proxy idle timers: stop reused sockets from failing the next request
- Forwarded client IP: trust a hop, not a request header
TLS and address families
Prove name selection and both IP families on the user path.
- TLS name selection: test SNI, certificate, and HTTP host together
- Dual-stack rollout: prove IPv4 and IPv6 reach the same service contract
Transport fallback and source metadata
Test UDP failure and isolate PROXY-aware listeners.
- HTTP/3 rollout: keep a tested TCP route when UDP fails
- PROXY protocol: accept client metadata only from the intended load balancer
Candidate and impact gates
Test the merged revision and fall back when selective coverage is uncertain.
- Merge queue checks: test the combined commit that will land
- Selective CI: default to full verification when impact is uncertain
Reliable test evidence
Keep flaky tests visible and isolate mutable data per run.
- Flaky tests: quarantine one failure mode with an owner and expiry
- CI test data: give each run an isolated ownership scope
Declared inputs and complete platforms
Control test dependencies and require every supported matrix row.
- Hermetic CI tests: declare every service and input the check consumes
- CI matrix gates: distinguish skipped, canceled, experimental, and passed
Artifact bytes and repeatability
Bind tests to the deployed digest and investigate unequal builds.
- Tested artifact identity: deploy the bytes that passed
- Reproducible builds: investigate why equal inputs produce different bytes
Writer and placement rules
Keep one writer and bind zonal storage with the consuming Pod.
- ReadWriteOncePod: enforce one Kubernetes writer for one claim
- Delayed volume binding: choose storage topology with the first Pod
Capacity and deletion paths
Check usable space after growth and trace the reclaim action before claim deletion.
- PVC expansion: verify both backing volume and usable filesystem
- PV reclaim policy: trace the real asset before deleting a claim
Ordinal and mount lifetime
Preserve StatefulSet claims intentionally and measure ownership work at startup.
- StatefulSet PVC retention: separate scale-down from deletion
- Volume ownership changes: keep fsGroup from turning recovery into a long mount
Device and node boundaries
Give raw block devices an application owner and plan for local-node loss.
- Raw block PVCs: make the application's formatting responsibility explicit
- Local PVs: plan for node loss as a data-location failure
Configuration propagation
Measure projected files, application reloads, and release-bound immutable generations.
- ConfigMap projection: prove when a running process sees a new value
- Immutable configuration: bind each release to one named generation
Startup and generation safety
Reject missing keys and mismatched policy-credential pairs before readiness.
- Required configuration keys: fail startup before accepting traffic
- Configuration pairs: prevent mixed policy and credential generations
Secret freshness and access
Observe issuer-to-process lag and workload creation as an indirect Secret permission.
- External secret sync: treat freshness as a monitored runtime contract
- Secret access: audit workload creation as an indirect read permission
Encryption and rotating identity
Keep decrypt keys available through restore and reopen projected token files.
- Kubernetes Secret encryption: rotate keys without losing restore access
- Projected identity tokens: reopen the file and verify its audience
Admission and syscall boundaries
Stage policy enforcement and verify runtime syscall filtering on each node pool.
- Pod Security Admission: stage warnings before a namespace denies Pods
- Seccomp RuntimeDefault: test syscall behavior across node runtimes
Node and process identity
Prove AppArmor profile availability and the actual process group set.
- AppArmor profiles: match Pod placement to actual node enforcement
- Supplemental groups: remove unexpected access inherited from an image
Filesystem and user mapping
Bound writable scratch and test user-namespace storage compatibility.
- Read-only root filesystems: inventory every required write path
- Pod user namespaces: verify host mapping and volume compatibility
Sandbox and local capacity
Price runtime overhead and include all local disk consumers in placement decisions.
- RuntimeClass: budget the real cost of a stronger sandbox
- Local ephemeral storage: account for logs, writable layers, and emptyDir
Allocation and request feedback
Reconcile the provider bill and test how request changes affect scheduling and HPA behavior.
- Shared cluster cost: reconcile service allocation with the provider bill
- CPU request rightsizing: account for the HPA feedback loop
Interruptions and commitments
Price replay work and distinguish discounted usage from guaranteed capacity.
- Interruptible compute: price recovery work, not only cheap node hours
- Compute commitments: separate billing coverage from machine availability
Network path cost
Measure cross-zone bytes and compare gateway and private-endpoint routes.
- Cross-zone data paths: measure bytes before changing placement
- NAT and private endpoints: compare the complete route and failure domain
Guardrails and consolidation
Use fast operational limits and prove a node can be removed without exhausting recovery headroom.
- Cost alerts: account for billing delay before a runaway resource spreads
- Node consolidation: calculate the capacity needed to evict safely
Version and lifecycle recovery
Find the exact retained version and test when lifecycle policy makes it unreachable or deletes it.
- Object version recovery: distinguish a delete marker from lost bytes
- Object lifecycle rules: model current, noncurrent, and archive states
Transfer and integrity
Clean up incomplete parts and verify copy bytes with an explicit digest contract.
- Multipart uploads: bound abandoned parts and retry cost
- Object integrity: verify bytes without trusting an ETag shortcut
Events and access
Process object notifications safely and constrain signed download capabilities.
- Object events: survive duplicate delivery and stale notifications
- Presigned object access: bound authority and expiry
Protection and reconciliation
Test held-version decryption and compare delayed inventory with live state.
- Locked objects: test retention, legal hold, and decryption together
- Object inventory: reconcile delayed snapshots with live decisions
Policy and role delegation
Trace authorization through policy layers and test constrained role creation.
- Cloud policy decisions: trace every authorization layer
- Delegated role creation: constrain both the new role and its use
Attributes and service trust
Protect access-control tags and bind service principals to intended sources.
- Tag-based access: defend the attribute write path
- Service-principal trust: bind the caller to the intended source
Session and key authority
Preserve temporary-session lineage and retire key grants safely.
- Temporary sessions: preserve caller lineage through role chains
- Encryption-key grants: inventory and retire temporary authority
Audit coverage and integrity
Prove which operations are logged and validate the delivered evidence chain.
- Audit trails: prove which data-plane actions are recorded
- Audit log integrity: verify delivery and digest continuity
Inventory scope and identity
Describe shipped components and bind the inventory to immutable release bytes.
- SBOM scope: distinguish build inputs from shipped components
- SBOM binding: keep the inventory attached to the tested digest
Advisory and exploitability decisions
Match advisories to actual components and keep not-affected claims reviewable.
- Package identity: verify advisory matches before changing a release
- VEX decisions: bind a not-affected claim to evidence and expiry
Build claims and admission
Approve the builder and exact subject before a workload starts.
- Build provenance: verify who asserted how the artifact was made
- Admission verification: compare the running image with approved evidence
Refresh and retention
Rebuild inherited packages and preserve evidence through registry cleanup.
- Base-image refresh: rebuild and retest when inherited bytes change
- Registry evidence retention: keep referrers with the release digest
Recovery clock and standby gates
Time the user journey and verify that the standby can serve it.
- Recovery-time budgets: measure every stage of service return
- Standby admission: prove the recovery region can accept real work
Declaration and ordered promotion
Give operators an authority path and fence writes before routing changes.
- Disaster declaration: define authority, scope, and stop conditions
- Failover ordering: fence, promote, validate, then move clients
Lost work and mixed clients
Reconcile acknowledged work and handle old connections during cutover.
- Write-gap ledgers: account for acknowledged work that never reached the standby
- Cutover drains: handle old connections and ambiguous client retries
Failback and evidence
Rejoin the former primary safely and reset recovery protection after a drill.
- Failback: rebuild the former primary before returning write authority
- Recovery drills: retain evidence and reset the next line of defense
GitOps render identity
Pin all render inputs and review complete environment output.
- GitOps render locks: identify every input behind an applied manifest
- Kustomize overlays: review the complete environment diff
Ownership and deletion
Give each resource a controller of record and preview pruning.
- GitOps ownership transfer: keep one reconciler authoritative per resource
- GitOps pruning: preview deletions as a separate release action
Truth of release status
Separate source, sync, health, and customer success signals.
- GitOps verification: separate source sync, resource health, and user success
- GitOps source outages: distinguish last-good operation from fresh deployment
Rollback and fleet control
Return intent and controller authority together, then limit target cohorts.
- GitOps rollback: return desired state and controller authority together
- Fleet promotion: bound the number of clusters changed at once
TLS identity and chain checks
Verify certificate paths and names from the real client matrix.
- TLS chain delivery: test the clients that actually connect
- TLS names: verify SNI selection and peer identity independently
Trust and serving rotation
Overlap private trust and prove processes serve the new certificate.
- Private CA rotation: overlap trust before changing issuers
- Certificate reloads: distinguish file delivery from active TLS state
Issuance and revocation
Test challenge paths, retry budgets, and actual client rejection.
- ACME issuance: test challenge reachability without burning production orders
- Certificate revocation: test the client behavior and status-service dependency
Compromise and workload identity
Replace exposed keys and constrain mTLS identities during migration.
- TLS key compromise: replace authority and end stale sessions
- mTLS workload identities: test both authentication and authorization on rotation
Platform ownership evidence
Verify reachable owners and compare declared dependencies with runtime evidence.
- Service catalog ownership: prove that the named team can respond
- Service catalog dependencies: distinguish declared edges from observed calls
Self-service resource lifecycle
Validate the request contract and recover from partial provisioning.
- Self-service provisioning: make the request contract explicit
- Provisioning state machines: handle partial success and retries
Template change boundaries
Upgrade generated consumers and restrict template execution privileges.
- Golden path templates: design upgrades after code generation
- Platform template execution: restrict parameters, actions, and credentials
Platform rollout and outcomes
Control shared API revisions and measure usable consumer outcomes.
- Platform API revisions: control changes to managed resources
- Platform product signals: measure usable outcomes and support load
Memory failure attribution
Separate limit kills, reclaim stalls, and retained memory before changing a budget.
- Container OOM attribution: separate a limit kill from node eviction
- JVM container memory: leave room beyond the Java heap
- Cgroup memory signals: read reclaim before the OOM counter
- Memory growth: distinguish retained data from useful cache
Memory placement and scratch space
Account for tmpfs files and replacement Pods as real memory consumers.
- Memory-backed emptyDir: budget file bytes as container memory
- Memory-safe rollouts: reserve space for old and new Pods
Serverless capacity and latency
Budget invocation time, downstream concurrency, and environment startup.
- Serverless invocation budgets: separate CPU, elapsed time, and I/O
- Serverless concurrency: cap the function before the database fails
- Serverless cold starts: measure initialization against the user path
Serverless work and release safety
Make redelivery, schedules, and version rollback explicit operating contracts.
- Serverless event idempotency: commit the effect and receipt together
- Scheduled functions: control overlap, catch-up, and missed runs
- Serverless rollback: restore code traffic without assuming data rolled back
Publishing and CMS state
Connect build identity, editorial status, and reliable change delivery.
- Web publishing evidence: prove the build, deployment, and live route separately
- Headless CMS publishing: define which records become public routes
- CMS webhooks: handle duplicate, reordered, and missing publication events
Public delivery and recovery
Bound staleness, isolate previews, and verify the page graph after rollback.
- Published content caches: bound the stale window and purge the right key
- Preview and production publishing: keep content and configuration separate
- Web release rollback: check the page graph and preserve content state
Browser policy and response coverage
Roll out CSP and verify headers on each response-producing path.
- Content Security Policy rollout: turn observed violations into an enforceable rule
- Security headers: verify static, server-rendered, and error responses independently
Origin, transport, and edge rules
Keep credentialed APIs, host transport, and WAF changes inside measured boundaries.
- Credentialed CORS: bind allowed origins to the response cache key
- HSTS rollout: inventory every hostname before extending transport policy
- WAF rule rollout: measure false positives before blocking editorial traffic
Delivery throughput evidence
Count active production changes and trace committed work to its first live artifact.
- Deployment frequency: count activated production changes from an event ledger
- Change lead time: join committed work to its first live artifact
Failure and recovery evidence
Attribute interventions, measure restored user paths, and classify repair releases.
- Change fail rate: attribute immediate intervention to a production deployment
- Failed deployment recovery time: keep impact, detection, and restoration clocks distinct
- Deployment rework: identify unplanned repair releases without hiding planned work
Kafka write and replay boundaries
Connect producer acknowledgments, retry behavior, and retained-log recovery.
- Kafka write durability: align acknowledgments with the in-sync replica floor
- Kafka producer retries: preserve partition order without claiming end-to-end exactly-once
- Kafka retention: size the replay window against the longest recovery path
Kafka state and consumer recovery
Rebuild keyed state, reset offsets safely, and plan partition capacity.
- Kafka log compaction: rebuild state without resurrecting deleted keys
- Kafka consumer offset recovery: preview every reset before changing a group
- Kafka partition keys: balance throughput without breaking per-entity order
RabbitMQ publish and consume safety
Separate routing, broker replication, and completed consumer effects.
- RabbitMQ publishing: distinguish broker confirmation from queue routing
- RabbitMQ quorum queues: plan node maintenance around majority availability
- RabbitMQ consumers: couple manual acknowledgments to a bounded prefetch window
RabbitMQ failure and migration
Repair poison messages, bound expiry, and migrate queue policy safely.
- RabbitMQ dead lettering: make poison-message transfer observable and recoverable
- RabbitMQ TTL: separate message expiry from immediate storage reclamation
- RabbitMQ queue policies: migrate immutable queue type without losing the handoff
Redis state and failover
Compare persistence, memory pressure, and asynchronous promotion.
- Redis persistence: state the recoverable write window before choosing AOF or snapshots
- Redis maxmemory: choose eviction behavior by the meaning of each key
- Redis Sentinel failover: reconcile acknowledged writes across a new primary
Redis routing and recovery
Test slot moves, pending stream work, and ephemeral notification gaps.
- Redis Cluster resharding: require redirect-aware clients and compatible key placement
- Redis Streams recovery: reconcile pending entries before trimming the log
- Redis keyspace notifications: keep correctness outside an ephemeral event channel
systemd service lifecycle
Test startup dependencies, crash-loop limits, and credential access.
- systemd dependencies: separate unit ordering from application readiness
- systemd restarts: bound crash loops and preserve evidence before resetting failures
- systemd service credentials: pass files through a narrow runtime boundary
systemd evidence and triggers
Retain logs, recover missed schedule work, and verify socket activation.
- systemd journals: retain failure evidence without exhausting the host disk
- systemd timers: treat missed-run catch-up as a trigger, not a work ledger
- systemd socket activation: preserve the listener without hiding service failure
Linux host change safety
Stage kernel reboots, SSH identity, and expiring login authority.
- Linux kernel rollouts: prove the running kernel after each reboot cohort
- SSH host-key rotation: change server identity without teaching clients to ignore warnings
- SSH user certificates: bind host login to an expiry and a narrow principal
Linux host recovery boundaries
Constrain sudo, stage firewall rules, and rehearse encrypted-disk unlock.
- sudo policy: authorize an exact maintenance action, not a path to a shell
- nftables rollouts: apply a checked ruleset with an access recovery timer
- LUKS recovery drills: preserve independent unlock and header recovery paths
Linux storage pressure and repair
Attribute full disks, grow layered filesystems, and contain read-only failures.
- Linux disk pressure: explain missing space before deleting application data
- Linux filesystem expansion: follow the block device to the mounted filesystem
- Linux read-only filesystem incidents: preserve evidence before repair
Linux storage continuity
Verify mounts, rebuild degraded arrays, and isolate I/O latency.
