A startup probe gives a slow-starting container time to initialize before liveness and readiness take over. A readiness probe decides whether a Pod should receive service traffic. A liveness probe decides whether Kubernetes should restart a container. These roles should use different failure criteria when the application can be alive but temporarily unable to serve.
Kubernetes probes: startup, readiness, and liveness
Operational decision
A search API loads an index before accepting traffic. Its startup endpoint fails until initialization finishes; readiness checks whether it can serve the current request class; liveness checks for a stuck process that a restart could plausibly repair. Do not make liveness depend on a shared database outage, because that can restart every replica during the same downstream failure. The fragment shows probe configuration for a container exposing port 8080. The handlers and timing need load tests: a probe that times out only during normal peak CPU can turn a healthy service into a restart loop. Watch startup duration and restart counts after each release, and verify readiness goes false during a deliberate shutdown before the process exits.
startupProbe:
httpGet: {path: /health/startup, port: 8080}
periodSeconds: 3
failureThreshold: 40
readinessProbe:
httpGet: {path: /health/ready, port: 8080}
periodSeconds: 5
livenessProbe:
httpGet: {path: /health/live, port: 8080}
periodSeconds: 10Cost and verification
The startup settings allow roughly two minutes for initialization before failure, aside from request timeouts and scheduling effects. Frequent probes create network and handler work; keep endpoints cheap and bounded. Readiness failure removes the Pod from normal service routing but does not restart it, while liveness failure can trigger a restart. A restart may make a transient dependency outage worse. Test thresholds under realistic CPU pressure and during termination, not only on a quiet laptop.
Common Mistakes
- Do not use the same dependency-heavy check for liveness and readiness.
- Do not set startup thresholds below ordinary cold-start time.
- Do not assume readiness catches incorrect business responses.
Connected lessons
- DevOps Tutorial
- Kubernetes Deployment: rolling update capacity
- Observability: join metrics, logs, and traces
- Incident response: contain impact, then learn
