Postgres included. No credit card.Start free
Guides · 3 min read

Readiness vs Liveness: Health Checks That Actually Help

Know which failures should stop traffic and which failures should restart a process.

Three questions, three signals

Readiness asks whether this instance should receive traffic now. Liveness asks whether restarting this process might recover it. Startup asks whether initial work has finished enough for the other probes to mean anything. The distinction matters because each failure causes a different action.

In Kubernetes, failed readiness removes a Pod from Service traffic. Failed liveness restarts its container. A startup probe delays liveness and readiness until startup succeeds. Other platforms use different names, but the traffic-versus-restart distinction is still useful.

A dependency outage is usually not a liveness failure

Suppose PostgreSQL is unavailable for thirty seconds. The API cannot serve data-backed requests, so readiness may fail. Restarting every API replica does not repair PostgreSQL; it creates connection storms when the database returns. Liveness should usually remain healthy unless the process itself is stuck.

yaml
readinessProbe:
  httpGet: { path: /health/ready, port: 3000 }
  periodSeconds: 5
livenessProbe:
  httpGet: { path: /health/live, port: 3000 }
  periodSeconds: 15
startupProbe:
  httpGet: { path: /health/live, port: 3000 }
  failureThreshold: 30
  periodSeconds: 2

These intervals are illustrative. Set failure thresholds from measured startup and recovery times. A readiness endpoint should not perform an expensive query every few seconds across hundreds of replicas; use a bounded check or cached dependency state.

Design endpoints around actions

/health/live can return 200 if the event loop is responsive and the process is not in an unrecoverable state. /health/ready can require completed initialization and essential dependencies. Do not put secrets, version internals, or full stack traces in a public health response.

Test three situations separately: slow cold start, temporary dependency loss, and an actual hung process. If a probe produces the wrong action in any case, change the endpoint or threshold before relying on it for automated recovery.

Probe thresholds are part of capacity planning

A readiness probe every five seconds with three failures gives roughly fifteen seconds before an unhealthy instance stops receiving traffic, plus routing propagation. That may be too slow for a crash and too aggressive for a brief database hiccup. Measure the failure pattern and adjust thresholds with the user-visible error budget in mind.

A liveness probe that restarts on a shared dependency failure can create a cascading incident. Every replica restarts, reloads caches, and reconnects simultaneously. The shared dependency receives a sudden burst exactly when it is weakest. Keeping liveness independent of shared dependencies reduces this risk; readiness can still prevent bad traffic.

Health checks also need a startup story. If a service takes ninety seconds to load a model, a liveness check after ten seconds can prevent it from ever becoming healthy. A startup probe gives the process a measured initialization window, after which ordinary readiness and liveness rules apply.

Further reading

Kubernetes probe guidance describes semantics and examples. Kubernetes Pod lifecycle explains readiness conditions and traffic behavior.