A load test applies a controlled workload and observes throughput, latency, errors, and resource saturation. Its value is in locating the limiting boundary, not in producing the largest requests-per-second number. A test that ignores request mix, data size, cache state, or downstream quotas can overstate production capacity.
Capacity and load tests: identify the next bottleneck
Operational decision
For a parcel-tracking API, replay a mix of lookup and update operations against synthetic records that resemble production sizes. Increase concurrency in steps, hold each step long enough to observe steady behavior, and record p95 latency, error ratio, CPU throttling, database connections, and queue age. Stop when an agreed safety threshold is reached; never run a destructive peak test against production without a controlled window and ownership. Use the small plan below to make assumptions visible. If a new replica improves throughput but doubles database connections, size the shared database before enabling autoscaling. Compare cold-cache and warm-cache runs, and include a rollout period when old and new Pods overlap. The observed bottleneck determines the next change: increasing replicas will not fix a serialized database lock.
service: parcel-tracking
trafficMix:
lookupPercent: 73
updatePercent: 27
steps:
- {clients: 35, holdMinutes: 9}
- {clients: 85, holdMinutes: 9}
- {clients: 155, holdMinutes: 9}
stopAt:
p95LatencyMs: 620
serverErrorPercent: 1.2Cost and verification
A realistic test needs isolated infrastructure and representative data, which cost time and compute. Short tests can miss garbage collection, cache eviction, connection leaks, and autoscaler delay. The numeric limits above are exercise values, not general service targets. Record test image digest, environment size, and data profile so two runs can be compared. Watch client-side saturation too; a load generator that cannot sustain its planned rate may make the service appear healthier than it is.
Common Mistakes
- Do not compare runs with different data profiles as if they were identical.
- Do not increase replicas before checking the downstream limit.
- Do not let the load generator become the hidden bottleneck.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes requests and limits: schedule for real load
- Horizontal autoscaling: choose a signal tied to demand
- SLOs and error budgets: turn reliability into a decision
Advanced follow-up
Advanced follow-up
Advanced follow-up
Memory failure follow-up
- Memory growth: distinguish retained data from useful cache
- Memory-safe rollouts: reserve space for old and new Pods
