A load test asks how a system behaves under a stated arrival pattern, not how fast one isolated request can complete. Describe the journey mix, active users or arrival rate, test data, warm-up, and duration. Record latency distributions, error rate, saturation signals, and the operation that slows first. A capacity budget is a decision boundary agreed before running the test: for example, a search route may need to serve 47 requests per second for several minutes while preserving a selected tail-latency and error target. The numbers are product-specific hypotheses to verify, not universal standards.
Load Tests and Capacity Budgets
Working case
After search launches, the case service sees bursts from shift changes. A test that repeatedly requests one cached case detail page reports a reassuring average. Real users search, filter, open cases, and save notes; those paths compete for database connections and index capacity. Build a mix that reflects these actions, seed enough cases to make the search index meaningful, and ramp arrivals to the expected burst. Watch high-percentile search latency, write failure rate, queue lag, CPU, and connection use. Stop when errors rise or the agreed budget is breached.
Implementation boundary
function withinCapacityBudget(result) {
return result.requestsPerSecond >= 47 && result.p95Milliseconds <= 420 &&
result.errorFraction < 0.01;
}
console.log(withinCapacityBudget({ requestsPerSecond: 49, p95Milliseconds: 390, errorFraction: 0.004 }));
// Output: trueThe function evaluates one report against an illustrative budget; a real test must measure the report with an instrumented runner and a stable environment. Separate offered load from completed throughput so overload cannot appear successful because the service simply drops work. Use open arrival models when the question is an external request rate, and user-loop models when the question is a population performing tasks with think time. Tag search, detail, and write operations so one cheap route cannot hide a slow critical route in the aggregate.
Cost and boundaries
A serious run consumes compute, test data, and engineer analysis time. Small smoke loads can run in a release pipeline; longer capacity studies belong in a controlled environment that resembles production enough to expose shared bottlenecks. Ramp slowly to distinguish steady behavior from cold starts, then include a short burst and recovery period. Higher concurrency may increase queueing faster than throughput. The test must not use private production records or send artificial writes into a live customer environment without a planned boundary.
Failure trace
A report checks mean latency only. Most search requests complete quickly, but one in twenty waits several seconds after the connection pool saturates. The average still looks acceptable, and the release proceeds. Add a tail-latency threshold and inspect pool wait time, database plans, and index lag under the same request mix. Another common error is a closed-loop script that slows its own request generation as the server slows, masking overload. Compare offered and completed rates and rerun after a targeted change.
Verification
- Measure offered and completed traffic with per-operation latency and error rates.
- Apply a predeclared tail-latency and error budget during sustained load and burst recovery.
- Inspect resource saturation and indexing lag before attributing a slow route to application code.
Practice drill
Seed a nonprivate case dataset with varied titles and permissions. Define a mixed workload of search, detail, and note writes, then state throughput, tail latency, and error limits before the run. Ramp to 47 requests per second, hold, spike briefly, and observe recovery. Capture per-operation results and resource saturation. Repeat after one index or pool change and compare the same workload. If the environment differs from production, write down the limits of the conclusion.
Decision note
Tie capacity claims to a measured workload and explicit thresholds; a fast average from one cached route is not release evidence.
Common Mistakes
- Using only mean latency as a release gate.
- Testing one cached request instead of the user journey mix.
- Letting a slowed closed-loop client silently reduce offered traffic.
Connected lessons
Quality and Capacity Engineering; Test Boundaries and Evidence Selection; Deterministic Test Doubles and Fault Injection; Accessibility and Visual Regression Checks; Page performance: budget the critical path and reserve layout space; Indexes and Query Plans for Case Feeds; Observability: connect user failure to a safe request trace.
Apply and check
Build Project: release evidence and capacity and review Web Development: search and quality contracts quiz.
Further connections
GraphQL Operation Cost and Admission.
Further connections
Connection Pool Budgets and Queue Admission.
