A shared NAT gateway translates private outbound connections into a finite set of address and port combinations. Many short connections to the same destination can exhaust available translations even when the application Pods, DNS, and remote API are healthy. The resulting connection errors can appear in several unrelated services at once because their egress path is shared.
NAT port pressure: find the shared outbound ceiling
Operational decision
A receipt service and an alert worker both start timing out against one payment endpoint after a rollout triples worker count. First group failures by destination address and port, source subnet, and gateway; check gateway port-allocation errors, active connections, and dropped packets over the same minutes. The read-only sample fetches recent gateway metrics in an approved account. Compare the trend with connection creation in clients. Restore bounded connection pooling and keep-alive, limit concurrent dials, and test a small traffic slice before increasing gateway address capacity. If a provider offers private connectivity, evaluate its cost and routing separately. An idle connection may be removed by network infrastructure; a pool must validate or replace it rather than treating any old socket as usable. Confirm successful application transactions after the change, not merely a falling gateway counter.
aws cloudwatch list-metrics --namespace AWS/NATGateway --metric-name ErrorPortAllocation --dimensions Name=NatGatewayId,Value=nat-0123456789abcdef0
aws cloudwatch list-metrics --namespace AWS/NATGateway --metric-name ActiveConnectionCount --dimensions Name=NatGatewayId,Value=nat-0123456789abcdef0Cost and verification
Connection reuse reduces port churn and handshake cost but holds sockets and application memory. More gateway addresses or separate egress paths add spend and can hide a client-side connection leak. Metric discovery does not return data points; graph the metric with the incident time window before deciding. Keep a per-destination connection budget and a test for idle-socket recovery. Scaling callers without checking the shared NAT path can make the outage wider.
Common Mistakes
- Do not read a healthy Pod network interface as proof the shared gateway has ports.
- Do not multiply clients while connection creation is exhausting a shared destination path.
- Do not assume an idle pooled socket survives every gateway timeout.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Egress policy and DNS: restrict destinations without breaking name resolution
- File descriptor exhaustion: find the leak before raising the limit
- Retries and timeouts: bound the cost of a failed request
- Capacity and load tests: identify the next bottleneck
