Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pod priority: reserve recovery capacity without evicting the wrong work

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A higher-priority Pod that cannot schedule may cause lower-priority Pods to be terminated so it can fit. Priority affects scheduling decisions and node-pressure eviction, but it does not create CPU, addresses, or attached-volume slots. A disruption budget influences some scheduler choices but should not be treated as an absolute shield against every form of disruption. A badly scoped high-priority class can displace important work across a shared cluster.

Operational decision

A payment gateway needs capacity during node failure while a batch document renderer can pause. Define a limited priority class for the gateway and restrict which teams may use it. The sample is a reviewable class manifest for a disposable cluster; admission and quota controls must separately prevent arbitrary use. Simulate a full node pool and create one replacement gateway Pod. Record which lower-priority Pods are selected, whether their shutdown preserves work, and whether the gateway becomes ready within its objective. Then repeat when a Pod requires an unavailable address or volume slot: preemption should not be presented as a cure for that different limit. Leave a documented batch replay path and alert on sustained Pending gateway Pods, preempted batch work, and unavailable payment replicas. A high priority for every workload collapses the ordering and makes the class useless.

yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: payment-recovery
value: 700000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: Reserved for payment gateway recovery Pods

Cost and verification

Reserved capacity costs money; relying only on preemption saves idle capacity but creates interruption and replay cost for lower-priority work. A grace period for victims can delay the replacement even after they are selected. High-priority usage needs a tenant quota and periodic audit. Track the replacement Pod's actual ready time and the business effect on displaced workloads, not only a scheduler event.

Common Mistakes

  • Do not assign the highest available priority to every production Pod.
  • Do not assume a disruption budget prevents every preemption or eviction.
  • Do not use priority as a substitute for IP, volume, or zone capacity.

Connected lessons

Practice and check

devops
operations
Storage details