A container resource request informs scheduling and is the denominator for CPU-utilization scaling. A limit constrains runtime use. CPU over a limit may be throttled; memory over a limit can end in an out-of-memory kill. Requests are not a promise that the application cannot use more, and a large request can prevent a Pod from being scheduled even when measured average use is low.
Kubernetes requests and limits: schedule for real load
Operational decision
For a document-index service, derive CPU and memory requests from representative load, then include cold starts, garbage-collection peaks, and the surge Pod needed by a rolling update. Observe throttling, memory working set, OOM events, and p95 request latency together. The fragment sets a 280 millicore CPU request and 640 MiB memory request. The CPU limit is higher to allow bursts, while the memory limit leaves headroom for a measured peak; these numbers are an exercise, not a recommendation for another service. Re-test after changing the runtime or payload size. If no node can fit the request, increasing replicas or enabling autoscaling will not solve placement until capacity changes.
resources:
requests:
cpu: 280m
memory: 640Mi
limits:
cpu: 900m
memory: 1180MiCost and verification
Higher requests reserve more scheduling capacity and can increase infrastructure cost. Low CPU limits can lengthen response time even when average usage looks safe; low memory limits can create restart loops. Omitting requests can make autoscaling by CPU utilization undefined for affected Pods. Compare costs against SLO impact under real peak traffic. A container can also exhaust ephemeral storage or file descriptors, which these CPU and memory settings do not cover.
Common Mistakes
- Do not copy requests from a quiet development environment.
- Do not confuse CPU throttling with memory OOM behavior.
- Do not add replicas when the scheduler has no room for them.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes Deployment: rolling update capacity
- Horizontal autoscaling: choose a signal tied to demand
- Capacity and load tests: identify the next bottleneck
Advanced follow-up
Advanced follow-up
Advanced follow-up
Advanced follow-up
Advanced follow-up
- RuntimeClass: budget the real cost of a stronger sandbox
- Local ephemeral storage: account for logs, writable layers, and emptyDir
