Zero-Downtime Deployments and Rollbacks Explained
Make old and new versions coexist, drain traffic safely, and recover without reversing data blindly.
“Zero downtime” is a property of the whole change
Replacing containers without a visible gap is only one part of a safe release. Users can still see errors if a new version expects a schema that old replicas cannot use, if sessions live in process memory, or if a dependent API changes incompatibly. The release unit includes code, data, configuration, and traffic routing.
During a rolling deployment, capacity is added or replaced in steps. A new instance should receive traffic only after readiness succeeds. An old instance should stop taking new traffic and finish active requests before exit. The platform needs enough spare capacity and a drain period for this handoff.
Make the previous version deployable
Store an immutable artifact reference for every release. If version N+1 increases errors, route traffic back to version N's artifact. Do not rebuild an old branch and assume it produces the same bytes. A deploy record should include the commit, artifact digest, configuration revision, and migration identifiers.
This is why destructive migrations are delayed. Container rollback cannot undo a dropped column or a payment already sent. For data changes, the recovery plan may be a forward fix rather than a database rollback.
Define a failure threshold before rollout
Readiness catches instances that never start. It does not catch a valid route that suddenly returns wrong prices. Add a smoke test for the most important user journey and watch error rate, latency, and business events during a release. Decide the observation window and rollback trigger before seeing the incident.
Amazon ECS, for example, can use a deployment circuit breaker to fail a rollout that does not reach steady state and optionally roll back to the last completed deployment. Application-level alarms can catch failures that a container health check misses.
Rehearse the recovery path
Run a deliberately broken release in staging. Verify that new instances fail readiness, traffic remains on the old version, and the deployment status reports failure. Then test a release that starts successfully but fails a business smoke check. The second case is where manual or metric-driven rollback procedures matter.
A rollout can be healthy and still be wrong
Suppose version N+1 answers /health/ready with 200 but serializes a checkout price incorrectly. The orchestrator sees a healthy task and may replace every old instance. A readiness probe cannot validate business correctness. Use a small post-deploy transaction, metric alarms, and, for high-risk changes, gradual traffic shifting or feature flags.
Rollback timing matters. If the bad release wrote data in a new format, version N may not understand it. The safest schema and data changes preserve old readers through the rollback window. If that is impossible, plan a forward repair and avoid claiming that one-click rollback solves the incident.
Sessions and background work need similar thought. An in-memory session disappears as traffic shifts between replicas. A queued job created by the new version may be picked up by an old worker after rollback. Version payloads or make consumers tolerant of the transition. Zero-downtime release design extends beyond the web process.
Further reading
ECS deployment circuit breaker documents automated rollback behavior. Kubernetes Deployments explains rolling update mechanics. See migration compatibility.