A capacity prompt should show each unit conversion and assumption before it recommends an instance count. Separate sustained demand, burst demand, measured sustainable capacity per worker, and reserve for an unavailable worker. Use a deterministic calculation for integer rounding. The number is a planning lower bound under the stated workload, not proof that latency, database capacity, or network limits will hold. Ask which resource saturates first and what load test would falsify the estimate. When the worker benchmark is missing, leave the denominator blank instead of dividing by a fabricated throughput.
Capacity prompts: check units, headroom, and a failed worker
Operational case
Receipt Ingest sees 18,600 requests per minute, or 310 per second. Applying the assumed 2.4 burst factor gives 744 requests per second. A load test under the same request mix measures 47 sustainable requests per second per healthy worker. The peak therefore needs 16 healthy workers because 15 provide only 705 per second, while 16 provide 752. To retain 16 after one worker fails, deploy 17. This ignores database, queue, and network limits until separate tests verify them. The estimate must be revised if the 2.4 factor or measured worker rate changes.
from math import ceil
normal_requests_per_minute = 18_600
planning_burst_factor = 2.4
measured_worker_rps = 47
failed_worker_reserve = 1
peak_rps = ceil(normal_requests_per_minute / 60 * planning_burst_factor)
healthy_workers = ceil(peak_rps / measured_worker_rps)
deployed_workers = healthy_workers + failed_worker_reserve
assert (deployed_workers - 1) * measured_worker_rps >= peak_rps
print(f"peak_rps={peak_rps} healthy_workers={healthy_workers} deployed_workers={deployed_workers}")Performance and operating cost
The arithmetic is O(1) time and space for fixed inputs; measurement and load testing dominate effort. At 47 requests per second, 16 healthy workers provide 752 per second, only eight above the planning peak, so the team should test burst duration, uneven routing, and downstream saturation before treating this as safe headroom. A queue can absorb a transient mismatch but adds storage, delay, and replay behavior. The prompt must state whether the arrival rate counts retries, because retries can raise offered load during failures.
Common Mistakes
- Do not round down a fractional worker requirement.
- Do not add a failed-worker reserve and then forget to evaluate capacity after that worker is gone.
- Do not claim end-to-end capacity from a worker-only benchmark.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Numeric prompts: let code calculate and the model explain
- Prompt budgets: trade output quality against cost and tail latency
- Aggregation prompts: pin the denominator and recompute the rate
- Architecture prompts: start with constraints and an owner
- Architecture decisions: compare real alternatives and consequences
- Architecture prompts: test the failure path and its signals
- Architecture release: require evidence, rollback, and ownership
- Project: review a receipt-ingest architecture decision
- Architecture-decision prompt checks
Continue with: Forecast prompts: keep scenarios separate from predictions.
