Deployment guide

Deploy FastAPI on a VPS: the complete production guide

Deploy on a VPS Cloud →

Tutorial

Deploy FastAPI on a VPS: the complete production guide

Development10 min read7 steps

FastAPI has established itself as the Python framework of choice for fast, typed APIs, thanks to its native support for async and OpenAPI documentation. Deploying it on your own VPS gives you control over the number of workers, WebSocket, and latency, with no ceiling imposed by a managed service. This guide covers the entire production setup: multi-stage containerization, database with Alembic migrations, automated CI/CD, and troubleshooting common errors.

Contents· Why Self-Host FastAPI on a VPS?1/9
  1. 01Why Self-Host FastAPI on a VPS?
  2. 02Concrete Benefits of a Self-Hosted FastAPI
  3. 03Hardware, Software Prerequisites and Versions
  4. 04Deploy FastAPI Step by Step
  5. 05Database and Migrations
  6. 06CI/CD and Automated Deployments
  7. 07Monitoring and Logs
  8. 08Troubleshooting Common Errors
  9. 09Deploy FastAPI in One Click from the ServOrbit Marketplace

Why Self-Host FastAPI on a VPS?

FastAPI draws all its performance from the asynchronous ASGI model: persistent WebSocket connections, response streaming, and concurrent calls to other services. Serverless platforms often cut long-lived connections and bill every request, which penalizes a high-throughput or real-time API. On a VPS, you run Uvicorn driven by Gunicorn with as many workers as cores, you keep connections open as long as needed, and you place Nginx at the front for TLS and buffering. You also control the connection pool to PostgreSQL and the Redis cache, two essential levers as load grows — the throughput actually sustained is measured against your own traffic.

Concrete Benefits of a Self-Hosted FastAPI

  • WebSockets and Server-Sent Events with no dropped connections or imposed timeout.
  • Fine-tuning of Uvicorn/Gunicorn workers according to the VPS's number of cores.
  • No cold start: the process stays up between requests.
  • Swagger and ReDoc documentation served internally, accessible or protected as you wish.
  • Direct integration with an asynchronous PostgreSQL pool (asyncpg) and Redis.
  • Fixed cost regardless of the number of API calls, ideal for a production backend.

Hardware, Software Prerequisites and Versions

A FastAPI API is lightweight: 1 vCPU and 1 GB of RAM are enough to start, but aim for 2 vCPU and 2 GB as soon as you add a database and a cache on the same VPS. For sustained load (>100 req/s), plan for at least 4 vCPU and 4 GB.

Recommended versions in 2025: Python 3.11+ (3.12 for maximum performance), FastAPI 0.115+ (full Pydantic v2 support and Python 3.10+ annotations), Uvicorn 0.30+ (event loop stability fixes). Install via Docker with the python:3.12-slim image, or from the official repository with pip install fastapi[standard]>=0.115 uvicorn[standard]>=0.30 gunicorn>=22.0.

On the infrastructure side: Docker Engine 24+, Docker Compose v2 (integrated plugin, docker compose), Nginx 1.24+ as reverse proxy, and a domain pointed to the VPS for the TLS certificate. Ubuntu 22.04/24.04 LTS is the recommended base; Debian 12 also works.

Deploy FastAPI Step by Step

  1. Initialize the VPS and the Application

    Connect via SSH, install Docker, clone your repository, and create a .env from the .env.example template provided with your project. This file must contain at minimum:

    DATABASE_URL=postgresql+asyncpg://user:password@db:5432/appdb
    REDIS_URL=redis://redis:6379/0
    SECRET_KEY=changeme-in-production
    CORS_ORIGINS=["https://yourdomain.com"]
    ENVIRONMENT=production

    Verify that your entry point exposes app = FastAPI() and that the application loads its configuration from environment variables (never hardcoded).

  2. Containerize with a Multi-Stage Dockerfile

    Multi-stage separates the build phase (dependency compilation, wheel generation) from the runtime image, reducing the final size by 60–80% and removing compilation tools from the attack surface:

    # ── Stage 1 : build ──────────────────────────────────────
    FROM python:3.12-slim AS builder
    WORKDIR /app
    COPY requirements.txt .
    RUN pip install --upgrade pip && \
        pip wheel --no-cache-dir --wheel-dir /wheels -r requirements.txt
    
    # ── Stage 2 : runtime ────────────────────────────────────
    FROM python:3.12-slim
    WORKDIR /app
    COPY --from=builder /wheels /wheels
    RUN pip install --no-cache-dir /wheels/* && rm -rf /wheels
    COPY . .
    EXPOSE 8000
    CMD ["gunicorn", "main:app", "-k", "uvicorn.workers.UvicornWorker",
         "--bind", "0.0.0.0:8000", "--workers", "4",
         "--timeout", "120", "--keep-alive", "5"]

    Add a /health endpoint to your application for healthchecks:

    @app.get("/health")
    async def health():
        return {"status": "ok"}
  3. Tune Gunicorn Workers Optimally

    The formula (2 × cores) + 1 is the starting point for CPU-bound workers. Since FastAPI is ASGI/async, each worker handles multiple requests concurrently via the event loop — for typical I/O-bound workloads (DB calls, external HTTP), cores + 1 to 2 × cores is more appropriate.

    Pass the value via an environment variable to adjust without rebuilding the image:

    WEB_CONCURRENCY=4  # in .env or docker-compose.yml
    CMD ["gunicorn", "main:app", "-k", "uvicorn.workers.UvicornWorker",
         "--bind", "0.0.0.0:8000",
         "--workers", "${WEB_CONCURRENCY:-4}"]

    Under high load, monitor CPU usage per worker (docker stats) and pending request count in Gunicorn logs (--access-logfile -) before increasing the worker count.

  4. Orchestrate Services with Docker Compose

    In docker-compose.yml, declare the api service, a PostgreSQL db service, and a redis service. Use depends_on with condition: service_healthy so the API waits for the database to be ready before starting:

    services:
      api:
        build: .
        env_file: .env
        ports: ["8000:8000"]
        depends_on:
          db:
            condition: service_healthy
          redis:
            condition: service_started
        healthcheck:
          test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
          interval: 30s
          retries: 3
    
      db:
        image: postgres:16-alpine
        environment:
          POSTGRES_USER: user
          POSTGRES_PASSWORD: password
          POSTGRES_DB: appdb
        volumes: ["pgdata:/var/lib/postgresql/data"]
        healthcheck:
          test: ["CMD", "pg_isready", "-U", "user"]
          interval: 10s
          retries: 5
    
      redis:
        image: redis:7-alpine
    
    volumes:
      pgdata:
  5. Place Nginx as a Reverse Proxy

    Configure a server block with proxy_pass http://api:8000 and, for WebSockets, add the necessary headers. Disable buffering for streaming if needed:

    server {
        listen 443 ssl;
        server_name api.yourdomain.com;
    
        location / {
            proxy_pass http://api:8000;
            proxy_set_header Host $host;
            proxy_set_header X-Real-IP $remote_addr;
            proxy_set_header Upgrade $http_upgrade;
            proxy_set_header Connection "upgrade";
            proxy_buffering off;
        }
    }
  6. Secure with a TLS Certificate

    Obtain a Let's Encrypt certificate via Certbot or a companion container, force HTTPS, and redirect HTTP. Remember to restrict access to /docs and /openapi.json in production if the API is not public.

  7. Launch and Test

    Start with docker compose up -d, check health via the /health endpoint you added, and test the interactive documentation. Monitor the Uvicorn logs to adjust the number of workers under real load.

For long-running tasks (sending emails, image processing, slow external calls), do not block the event loop: offload them to a background worker with BackgroundTasks for simple cases, or to Celery/ARQ with Redis for heavy processing. An API that responds quickly and delegates asynchronous work stays responsive even under high concurrency.

Database and Migrations

For FastAPI applications in production, asyncpg combined with SQLAlchemy 2.x (async mode) offers an excellent balance of performance and maintainability. Configure the connection pool according to expected load:

from sqlalchemy.ext.asyncio import create_async_engine, AsyncSession
from sqlalchemy.orm import sessionmaker

engine = create_async_engine(
    DATABASE_URL,
    pool_size=10,        # persistent connections
    max_overflow=20,     # additional connections under peak
    pool_timeout=30,     # seconds before error
    pool_recycle=1800,   # recycle after 30 min to avoid PostgreSQL timeouts
    echo=False,
)

AsyncSessionLocal = sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)

Inject the session into your routes via a FastAPI dependency:

async def get_db():
    async with AsyncSessionLocal() as session:
        yield session

For schema migrations, Alembic is the standard tool. Initialize it in your project (alembic init migrations) then configure alembic.ini to point to your DATABASE_URL. Essential commands:

# Generate a migration from your SQLAlchemy models
alembic revision --autogenerate -m "add_users_table"

# Apply pending migrations
alembic upgrade head

# Check current state
alembic current

In production, run alembic upgrade head before starting the API — add it as an init command in your docker-compose.yml or in a separate startup script. Never run migrations from multiple instances simultaneously.

CI/CD and Automated Deployments

A simple GitHub Actions pipeline is sufficient for most FastAPI projects: build the image, push to a registry, and deploy to the VPS via SSH with zero-downtime.

# .github/workflows/deploy.yml
name: Deploy FastAPI
on:
  push:
    branches: [main]

jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Build & push Docker image
        run: |
          docker build -t ghcr.io/${{ github.repository }}:${{ github.sha }} .
          echo ${{ secrets.GITHUB_TOKEN }} | docker login ghcr.io -u ${{ github.actor }} --password-stdin
          docker push ghcr.io/${{ github.repository }}:${{ github.sha }}

      - name: Deploy via SSH
        uses: appleboy/ssh-action@v1
        with:
          host: ${{ secrets.VPS_HOST }}
          username: deploy
          key: ${{ secrets.VPS_SSH_KEY }}
          script: |
            cd /opt/myapp
            export IMAGE_TAG=${{ github.sha }}
            docker compose pull api
            docker compose up -d --no-deps api
            docker compose exec api alembic upgrade head

The --no-deps flag on docker compose up restarts only the api container without touching db or redis, ensuring zero-downtime for third-party services. For more robust deployments with instant rollback, tag your images with the commit SHA and keep the last 3 versions locally.

Monitoring and Logs

Access Uvicorn logs directly via Docker:

# Real-time logs
docker compose logs -f api

# Last 100 lines
docker compose logs --tail=100 api

For structured JSON logs (indexable by an aggregator like Loki or Elasticsearch), replace the standard Python logger with loguru or structlog:

from loguru import logger
import sys

logger.remove()
logger.add(sys.stdout, format="{time} {level} {message}", serialize=True)

@app.middleware("http")
async def log_requests(request, call_next):
    logger.info("request", method=request.method, url=str(request.url))
    response = await call_next(request)
    logger.info("response", status=response.status_code)
    return response

To expose Prometheus metrics, add a /metrics endpoint with prometheus-fastapi-instrumentator:

from prometheus_fastapi_instrumentator import Instrumentator

Instrumentator().instrument(app).expose(app, endpoint="/metrics")

Then restrict /metrics to your internal network from Nginx (allow 127.0.0.1; deny all;) to avoid publicly exposing metrics.

Troubleshooting Common Errors

Worker timed out (pid XXXXX) — Gunicorn kills a worker that did not respond within the --timeout delay. Common causes: a synchronous route (def instead of async def) blocks the event loop, or a blocking operation (file read, network call) runs without asyncio.run_in_executor. Increasing --timeout masks the problem; the correct fix is to make the route async or offload blocking work.

Address already in use (port 8000) — A previous Gunicorn process is still active. In Docker: docker compose down then docker compose up -d forces container recreation and releases the port. Check with docker ps -a that no zombie container is running.

502 Bad Gateway after a deployment — Nginx forwarded the request to the API during its restart. Wait 10–15 seconds for the Docker healthcheck to confirm the new container is healthy before considering the deployment complete. Add proxy_read_timeout 60; in Nginx to absorb slow startups.

CORS in production: No 'Access-Control-Allow-Origin' header — FastAPI serves the origins declared in CORSMiddleware. Verify that CORS_ORIGINS in your .env contains exactly the frontend URL (protocol + domain + port if non-standard, no trailing slash). A value of ["*"] works in development but must never go to production.

DB connection refused at startup — The API starts before PostgreSQL is ready. depends_on with condition: service_healthy in Docker Compose solves this. In Kubernetes or deployments without Compose, implement a retry loop with exponential backoff in the application startup code.

Deploy FastAPI in One Click from the ServOrbit Marketplace

The ServOrbit Marketplace offers a preconfigured FastAPI Stack template: Python 3.12, Uvicorn, Gunicorn, Nginx, and PostgreSQL 16 are installed automatically on your VPS from the order form. No SSH session for the base setup — you arrive directly at the repository clone step and start Uvicorn. The marketplace page contains the exact specifications, use cases, and a first-login guide.

Deploy Your FastAPI Stack in Minutes

The ServOrbit FastAPI Stack template installs Python 3.12, Uvicorn, Gunicorn, Nginx, and PostgreSQL 16 on your VPS. Launch your API immediately.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab