A bulkhead reserves independent execution capacity for distinct dependency paths. Without one, a slow optional service can occupy every request thread or connection slot and block core transactions that do not need it. The isolation must include both active calls and waiting work; a limit on active calls with an unlimited pending queue merely moves the overload into memory and latency.
Dependency bulkheads: stop one slow path consuming every worker
Operational decision
A claims API calls a document renderer for previews and a ledger for settlement. Under renderer slowdown, keep settlement calls in a separate bounded pool with its own timeout and queue. The text contract sets test values for a disposable workload; tune them from actual request rate and latency distributions. Reject excess preview work with a clear status and retry guidance when safe, while keeping the ledger path inside its own capacity budget. Load the renderer until its queue fills and verify settlements still meet their latency target. Then load the ledger independently so the reserve is not merely unused capacity in the normal case. Watch active calls, queue depth, queue wait, rejected requests, and settlement success, grouped by dependency. Preserve end-to-end deadlines: work that waited too long in a local queue should not start after the caller has already given up.
Claims API concurrency partition
Document preview: 12 active, 24 waiting, 2-second call deadline
Ledger settlement: 8 active, 8 waiting, 4-second call deadline
Admission: reject when path-specific queue is full
Deadline: include local queue wait and remote call
Acceptance: renderer stall does not breach settlement objectiveCost and verification
Reserved capacity can sit idle and raise compute cost, but protects the core path during a dependency incident. Setting pools too small creates avoidable rejection; setting them too large can exhaust the remote service or database. A shared CPU quota, NAT gateway, or process file limit can still couple the paths even after thread pools are split. Test these shared limits as part of the isolation claim.
Common Mistakes
- Do not create separate active-call limits with one unlimited shared queue.
- Do not start remote work after its end-to-end deadline has expired.
- Do not claim isolation without testing shared CPU, network, and descriptor ceilings.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Overload shedding: refuse excess work before latency collapses
- Database pool pressure: bound waiting before the database collapses
- NAT port pressure: find the shared outbound ceiling
- Retries and timeouts: bound the cost of a failed request
