Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: harden a document worker without losing its recovery path

Last updated: 5 Oct 20269 min read
project
AdvancedBy AITrove Editorial

Use disposable namespaces and node pools with known kernel, runtime, AppArmor, and storage capabilities. Run a document-unpacking worker against synthetic archives, including one oversized file and one file intended to test a denied policy write. Record the exact manifest revision, node image, security profile hash, and active process identity before fault injection. Keep the worker out of production traffic until the full unpack, upload, restart, and rollback path is observed.

Prove the boundaries

Stage baseline enforce with restricted warn and audit, then server-side dry-run every workload template. Move to restricted enforcement only after a replacement Pod survives a drain. Test RuntimeDefault seccomp across two node-runtime builds. Install a custom AppArmor profile on only part of a disposable node pool and verify unsupported placement fails visibly; then distribute it and repeat. Compare process groups under Merge and Strict on supported nodes, using a decoy file to expose unintended access. Confirm the image root is read-only and that every required write lands in a bounded scratch mount.

Output
Document worker acceptance gates
Admission: replacement Pod passes reviewed policy version
Seccomp: supported render and unpack paths work on every node build
AppArmor: profile hash and denied write are verified
Groups: process lacks undeclared image groups
Root: required writes use bounded scratch, never the image layer
User namespace: host mapping and every volume mount are tested
RuntimeClass: handler exists and overhead fits surge capacity
Disk: logs, layers, and emptyDir stay inside node headroom

Inject failure and recover

Launch a user-namespace canary with compatible storage, then try a separately isolated incompatible-volume test and record the mount failure. Use an isolated RuntimeClass on prepared nodes; remove one eligible node and verify replacement headroom. Fill scratch, then grow logs and the writable layer independently to exercise local-storage accounting. Observe whether the Pod is evicted and whether the queue lease releases for replay. Roll back the hardening change through the reviewed manifest, preserving the denied-path audit evidence. Do not weaken every workload's policy because one worker has an undeclared file path.

Cost and verification

Report policy violations, denied syscalls and file accesses, node-specific start failures, initial process groups, eligible-node capacity, sandbox cold-start cost, and disk high-water marks. State which behavior was verified on an actual node and which was only parsed as configuration. A manifest accepted by the API server does not prove the runtime handler or security profile exists on the scheduled node.

Common Mistakes

  • Do not enforce a new policy before testing replacement Pods.
  • Do not assume a declared profile is present on every node.
  • Do not count only emptyDir when sizing local ephemeral storage.

Connected lessons

Memory failure follow-up

devops
project
Storage details