How to Deploy FastAPI Beyond the Development Server
Package a Python API with pinned dependencies, an ASGI server, health checks, and environment-based configuration.
FastAPI is the application, ASGI is the serving contract
FastAPI does not listen for network traffic by itself. An ASGI server such as Uvicorn imports the application object, accepts connections, and passes requests into it. That distinction explains why uvicorn main:app needs both the Python module name and the object name.
In a container, listen on 0.0.0.0 and the port expected by the platform. Run a production dependency install from a pinned file or lockfile. A local virtual environment on macOS does not prove Linux wheels or system libraries will be available in the build image.
This shell command works when the deployment start command is interpreted by a shell. In an exec-form Docker CMD, shell expansion does not occur; use a small entrypoint script or an explicit server command that reads configuration.
Choose workers from resource limits, not folklore
More workers can serve more concurrent CPU-bound requests, but every worker is a separate process with its own memory, connections, and startup work. Loading a large ML model in four workers may load it four times. In a container platform, scaling replicas and scaling workers are distinct choices.
Start with one process per container and measure latency, CPU, memory, and queueing. For I/O-heavy routes, async code helps only when called libraries are actually nonblocking. A synchronous database call inside an async def route can still block that worker's event loop.
Keep startup and health observable
Use application lifespan hooks for required initialization. Expose a cheap readiness route that fails until startup completes. Run schema migrations as a controlled release step, not in every worker import: multiple replicas racing to migrate can lock tables or produce duplicate side effects.
Long AI or import tasks should not hold an HTTP request open indefinitely. Queue the work, return a job identifier, and let a worker report completion. This isolates proxy timeouts from actual job execution.
A production smoke test
- Build from a clean checkout on the target OS or container base.
- Start with the exact production command and environment variable names.
- Request the OpenAPI route, a real business route, and readiness.
- Confirm CORS allows only intended browser origins.
- Inspect how the process exits on SIGTERM.
When a healthy process is still slow
FastAPI can accept asynchronous requests, but concurrency is not unlimited. A CPU-heavy model call occupies CPU even if its route is async. A synchronous client inside an async route can block the event loop. Measure where time is spent before adding workers: a slow database query, external API, or large response may be the actual bottleneck.
Suppose each worker loads a 1 GB model. Four workers in one container can need roughly four copies, plus Python and request overhead. A memory limit that worked with one worker may start killing the process after scaling workers. A separate model worker or job queue may be a better architecture than raising the web worker count.
If startup downloads a model or warms a cache, readiness must stay false until this work completes. Give the platform enough startup time, but do not hide a permanently failing initialization behind a very long grace period. Log the exact stage that failed and fail the process decisively when required assets cannot load.
Further reading
FastAPI deployment concepts explains processes, servers, and containers. FastAPI lifespan events covers application initialization and shutdown.