Deploy an Express API: Ports, Health Checks, and Shutdown
Turn a local Node API into a service that starts reliably and survives rolling deploys.
The process contract
An Express API becomes a deployable service when it obeys four contracts: start with one command, listen on the assigned port, report readiness, and stop accepting traffic on termination. Code that works at localhost:3000 can fail behind a container proxy if it binds only to loopback.
Read the port from the environment and bind to 0.0.0.0. Keep durable state outside the process. A container restart erases in-memory sessions, uploads, and pending jobs. Multiple replicas do not share memory, so such state also causes inconsistent behavior before a restart.
Readiness is different from process existence
A process can be running while its database pool is still connecting. A readiness endpoint should return success only when the service can handle its essential request path. Keep the check cheap and bounded; a slow probe can itself create load. A liveness endpoint should be narrower and detect a process that needs restarting, not every shared dependency outage.
This illustrates the lifecycle, not a complete health implementation. In a real app, set readiness after migrations or required connections are established. Add a bounded shutdown deadline for stuck requests and close database connections after the HTTP server drains.
Failure paths to exercise locally
Build a production artifact and run the same start command used by the host. Send a request to the readiness route, then send SIGTERM while a request is active. Verify that the process stops taking new requests and completes or times out existing work.
Test a missing required environment variable too. Fail fast with a clear error rather than starting a server that returns 500 on every request. Structured logs should include request IDs and errors, but never authorization headers or connection strings.
A shutdown sequence under real traffic
Imagine a rollout while a user is uploading a file. The platform sends SIGTERM to the old instance. If the process exits immediately, the upload fails. If it never exits, the rollout stalls or the platform eventually kills it. The application should first fail readiness, then stop accepting new connections, allow active work a bounded interval to finish, and finally close database or queue clients.
The sample code calls server.close(), which stops accepting new connections and waits for active requests. Add a timer that forces exit after the platform's grace period, and log whether it was needed. Keep long-running jobs out of the HTTP process so shutdown remains predictable.
There is another multi-replica trap: sticky behavior hidden in memory. If login sessions or rate limits live only in a JavaScript map, users may appear logged out or limits may reset when a request reaches another replica. Put shared state in an external store or use stateless signed sessions with appropriate key rotation.
Further reading
Node's `server.close()` behavior explains how the HTTP server stops accepting connections. Kubernetes probe semantics provides a useful model even when the runtime is not Kubernetes.