A policy exception permits a narrowly defined action that the ordinary deployment or admission rule would reject. It should not turn a cluster-wide rule off for convenience. The exception needs a bounded target, business reason, expiry, owner, approval record, and evidence that the baseline rule still denies all other workloads. An expiry timestamp in a ticket is only useful if a process acts on it.
Policy exceptions: make a temporary bypass expire and prove its scope
Operational decision
A document conversion worker needs a temporary write path while its volume driver is upgraded. Keep the admission rule in force for the namespace and define a workload-specific exception using the chosen policy engine's supported mechanism. The text record is a review template, not an admission manifest. Test one allowed admission for the named worker and one denied admission for a similar Pod. Apply a compensating filesystem restriction and monitor writes during the window. At expiry, remove the exception, rerun both tests, and check the running workload no longer depends on it. If the upgrade is incomplete, renew through a fresh risk decision rather than silently changing the date. Inventory exceptions from the policy system, not only from a spreadsheet, so an abandoned rule cannot survive indefinitely.
Temporary admission exception
Workload: document-converter in render namespace
Reason: volume-driver transition
Owner: render platform team
Expiry: 2026-10-18 18:00 UTC
Compensation: read-only root plus named writable mount
Proof: named Pod allowed; peer Pod denied
Removal: delete exception and repeat both admission testsCost and verification
Narrow exceptions demand review and test time but preserve the default control for other workloads. An open-ended exception saves immediate work while increasing exposure and making future audits harder. Policy-engine logs can show admissions, yet they may not reveal all runtime effects; pair admission checks with a workload behavior test. Track exception age and the count of workloads actually matching it, and page the owner before expiry when removal could break service.
Common Mistakes
- Do not disable a namespace-wide or cluster-wide rule for one workload.
- Do not rely on a ticket expiry without checking the live policy object.
- Do not remove an exception without verifying the repaired workload path.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes admission policy: reject an unsafe workload before scheduling
- Container runtime restrictions: remove privileges a service does not use
- Release evidence: tie one deployed digest to one approval decision
- Incident reviews: turn a timeline into tested corrective work
Practice and check
Cloud authority follow-up
- Tag-based access: defend the attribute write path
- Service-principal trust: bind the caller to the intended source
