Postgres included. No credit card.Start free
Deployment · 3 min read

Zero-Downtime Deployments and Rollbacks Explained

Make old and new versions coexist, drain traffic safely, and recover without reversing data blindly.

“Zero downtime” is a property of the whole change

Replacing containers without a visible gap is only one part of a safe release. Users can still see errors if a new version expects a schema that old replicas cannot use, if sessions live in process memory, or if a dependent API changes incompatibly. The release unit includes code, data, configuration, and traffic routing.

During a rolling deployment, capacity is added or replaced in steps. A new instance should receive traffic only after readiness succeeds. An old instance should stop taking new traffic and finish active requests before exit. The platform needs enough spare capacity and a drain period for this handoff.

Make the previous version deployable

Store an immutable artifact reference for every release. If version N+1 increases errors, route traffic back to version N's artifact. Do not rebuild an old branch and assume it produces the same bytes. A deploy record should include the commit, artifact digest, configuration revision, and migration identifiers.

text
release N → expand schema → release N+1 → verify → contract later
               ↘ if N+1 fails, release N still understands schema

This is why destructive migrations are delayed. Container rollback cannot undo a dropped column or a payment already sent. For data changes, the recovery plan may be a forward fix rather than a database rollback.

Define a failure threshold before rollout

Readiness catches instances that never start. It does not catch a valid route that suddenly returns wrong prices. Add a smoke test for the most important user journey and watch error rate, latency, and business events during a release. Decide the observation window and rollback trigger before seeing the incident.

Amazon ECS, for example, can use a deployment circuit breaker to fail a rollout that does not reach steady state and optionally roll back to the last completed deployment. Application-level alarms can catch failures that a container health check misses.

Rehearse the recovery path

Run a deliberately broken release in staging. Verify that new instances fail readiness, traffic remains on the old version, and the deployment status reports failure. Then test a release that starts successfully but fails a business smoke check. The second case is where manual or metric-driven rollback procedures matter.

A rollout can be healthy and still be wrong

Suppose version N+1 answers /health/ready with 200 but serializes a checkout price incorrectly. The orchestrator sees a healthy task and may replace every old instance. A readiness probe cannot validate business correctness. Use a small post-deploy transaction, metric alarms, and, for high-risk changes, gradual traffic shifting or feature flags.

Rollback timing matters. If the bad release wrote data in a new format, version N may not understand it. The safest schema and data changes preserve old readers through the rollback window. If that is impossible, plan a forward repair and avoid claiming that one-click rollback solves the incident.

Sessions and background work need similar thought. An in-memory session disappears as traffic shifts between replicas. A queued job created by the new version may be picked up by an old worker after rollback. Version payloads or make consumers tolerant of the transition. Zero-downtime release design extends beyond the web process.

Further reading

ECS deployment circuit breaker documents automated rollback behavior. Kubernetes Deployments explains rolling update mechanics. See migration compatibility.