Use disposable namespaces and node pools with known kernel, runtime, AppArmor, and storage capabilities. Run a document-unpacking worker against synthetic archives, including one oversized file and one file intended to test a denied policy write. Record the exact manifest revision, node image, security profile hash, and active process identity before fault injection. Keep the worker out of production traffic until the full unpack, upload, restart, and rollback path is observed.
Project: harden a document worker without losing its recovery path
Prove the boundaries
Stage baseline enforce with restricted warn and audit, then server-side dry-run every workload template. Move to restricted enforcement only after a replacement Pod survives a drain. Test RuntimeDefault seccomp across two node-runtime builds. Install a custom AppArmor profile on only part of a disposable node pool and verify unsupported placement fails visibly; then distribute it and repeat. Compare process groups under Merge and Strict on supported nodes, using a decoy file to expose unintended access. Confirm the image root is read-only and that every required write lands in a bounded scratch mount.
Document worker acceptance gates
Admission: replacement Pod passes reviewed policy version
Seccomp: supported render and unpack paths work on every node build
AppArmor: profile hash and denied write are verified
Groups: process lacks undeclared image groups
Root: required writes use bounded scratch, never the image layer
User namespace: host mapping and every volume mount are tested
RuntimeClass: handler exists and overhead fits surge capacity
Disk: logs, layers, and emptyDir stay inside node headroomInject failure and recover
Launch a user-namespace canary with compatible storage, then try a separately isolated incompatible-volume test and record the mount failure. Use an isolated RuntimeClass on prepared nodes; remove one eligible node and verify replacement headroom. Fill scratch, then grow logs and the writable layer independently to exercise local-storage accounting. Observe whether the Pod is evicted and whether the queue lease releases for replay. Roll back the hardening change through the reviewed manifest, preserving the denied-path audit evidence. Do not weaken every workload's policy because one worker has an undeclared file path.
Cost and verification
Report policy violations, denied syscalls and file accesses, node-specific start failures, initial process groups, eligible-node capacity, sandbox cold-start cost, and disk high-water marks. State which behavior was verified on an actual node and which was only parsed as configuration. A manifest accepted by the API server does not prove the runtime handler or security profile exists on the scheduled node.
Common Mistakes
- Do not enforce a new policy before testing replacement Pods.
- Do not assume a declared profile is present on every node.
- Do not count only emptyDir when sizing local ephemeral storage.
Connected lessons
- Pod Security Admission: stage warnings before a namespace denies Pods
- Seccomp RuntimeDefault: test syscall behavior across node runtimes
- AppArmor profiles: match Pod placement to actual node enforcement
- Supplemental groups: remove unexpected access inherited from an image
- Read-only root filesystems: inventory every required write path
- Pod user namespaces: verify host mapping and volume compatibility
- RuntimeClass: budget the real cost of a stronger sandbox
- Local ephemeral storage: account for logs, writable layers, and emptyDir
- DevOps projects
