A higher-priority Pod that cannot schedule may cause lower-priority Pods to be terminated so it can fit. Priority affects scheduling decisions and node-pressure eviction, but it does not create CPU, addresses, or attached-volume slots. A disruption budget influences some scheduler choices but should not be treated as an absolute shield against every form of disruption. A badly scoped high-priority class can displace important work across a shared cluster.
Pod priority: reserve recovery capacity without evicting the wrong work
Operational decision
A payment gateway needs capacity during node failure while a batch document renderer can pause. Define a limited priority class for the gateway and restrict which teams may use it. The sample is a reviewable class manifest for a disposable cluster; admission and quota controls must separately prevent arbitrary use. Simulate a full node pool and create one replacement gateway Pod. Record which lower-priority Pods are selected, whether their shutdown preserves work, and whether the gateway becomes ready within its objective. Then repeat when a Pod requires an unavailable address or volume slot: preemption should not be presented as a cure for that different limit. Leave a documented batch replay path and alert on sustained Pending gateway Pods, preempted batch work, and unavailable payment replicas. A high priority for every workload collapses the ordering and makes the class useless.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: payment-recovery
value: 700000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: Reserved for payment gateway recovery PodsCost and verification
Reserved capacity costs money; relying only on preemption saves idle capacity but creates interruption and replay cost for lower-priority work. A grace period for victims can delay the replacement even after they are selected. High-priority usage needs a tenant quota and periodic audit. Track the replacement Pod's actual ready time and the business effect on displaced workloads, not only a scheduler event.
Common Mistakes
- Do not assign the highest available priority to every production Pod.
- Do not assume a disruption budget prevents every preemption or eviction.
- Do not use priority as a substitute for IP, volume, or zone capacity.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Pod disruption budgets: make node drains measurable
- Namespace quotas: reserve room for a safe rollout
- Node autoscaling: make pending Pods schedulable before traffic rises
- Queue consumers: acknowledgement, idempotency, and backlog
