Why Self-Host FastAPI on a VPS?
FastAPI draws all its performance from the asynchronous ASGI model: persistent WebSocket connections, response streaming, and concurrent calls to other services. Serverless platforms often cut long-lived connections and bill every request, which penalizes a high-throughput or real-time API. On a VPS, you run Uvicorn driven by Gunicorn with as many workers as cores, you keep connections open as long as needed, and you place Nginx at the front for TLS and buffering. You also control the connection pool to PostgreSQL and the Redis cache, two essential levers as load grows — the throughput actually sustained is measured against your own traffic.
Concrete Benefits of a Self-Hosted FastAPI
- WebSockets and Server-Sent Events with no dropped connections or imposed timeout.
- Fine-tuning of Uvicorn/Gunicorn workers according to the VPS's number of cores.
- No cold start: the process stays up between requests.
- Swagger and ReDoc documentation served internally, accessible or protected as you wish.
- Direct integration with an asynchronous PostgreSQL pool (asyncpg) and Redis.
- Fixed cost regardless of the number of API calls, ideal for a production backend.
Hardware, Software Prerequisites and Versions
A FastAPI API is lightweight: 1 vCPU and 1 GB of RAM are enough to start, but aim for 2 vCPU and 2 GB as soon as you add a database and a cache on the same VPS. For sustained load (>100 req/s), plan for at least 4 vCPU and 4 GB.
Recommended versions in 2025: Python 3.11+ (3.12 for maximum performance), FastAPI 0.115+ (full Pydantic v2 support and Python 3.10+ annotations), Uvicorn 0.30+ (event loop stability fixes). Install via Docker with the python:3.12-slim image, or from the official repository with pip install fastapi[standard]>=0.115 uvicorn[standard]>=0.30 gunicorn>=22.0.
On the infrastructure side: Docker Engine 24+, Docker Compose v2 (integrated plugin, docker compose), Nginx 1.24+ as reverse proxy, and a domain pointed to the VPS for the TLS certificate. Ubuntu 22.04/24.04 LTS is the recommended base; Debian 12 also works.
Deploy FastAPI Step by Step
Initialize the VPS and the Application
Connect via SSH, install Docker, clone your repository, and create a
.envfrom the.env.exampletemplate provided with your project. This file must contain at minimum:DATABASE_URL=postgresql+asyncpg://user:password@db:5432/appdb REDIS_URL=redis://redis:6379/0 SECRET_KEY=changeme-in-production CORS_ORIGINS=["https://yourdomain.com"] ENVIRONMENT=productionVerify that your entry point exposes
app = FastAPI()and that the application loads its configuration from environment variables (never hardcoded).Containerize with a Multi-Stage Dockerfile
Multi-stage separates the build phase (dependency compilation, wheel generation) from the runtime image, reducing the final size by 60–80% and removing compilation tools from the attack surface:
# ── Stage 1 : build ────────────────────────────────────── FROM python:3.12-slim AS builder WORKDIR /app COPY requirements.txt . RUN pip install --upgrade pip && \ pip wheel --no-cache-dir --wheel-dir /wheels -r requirements.txt # ── Stage 2 : runtime ──────────────────────────────────── FROM python:3.12-slim WORKDIR /app COPY --from=builder /wheels /wheels RUN pip install --no-cache-dir /wheels/* && rm -rf /wheels COPY . . EXPOSE 8000 CMD ["gunicorn", "main:app", "-k", "uvicorn.workers.UvicornWorker", "--bind", "0.0.0.0:8000", "--workers", "4", "--timeout", "120", "--keep-alive", "5"]Add a
/healthendpoint to your application for healthchecks:@app.get("/health") async def health(): return {"status": "ok"}Tune Gunicorn Workers Optimally
The formula
(2 × cores) + 1is the starting point for CPU-bound workers. Since FastAPI is ASGI/async, each worker handles multiple requests concurrently via the event loop — for typical I/O-bound workloads (DB calls, external HTTP),cores + 1to2 × coresis more appropriate.Pass the value via an environment variable to adjust without rebuilding the image:
WEB_CONCURRENCY=4 # in .env or docker-compose.ymlCMD ["gunicorn", "main:app", "-k", "uvicorn.workers.UvicornWorker", "--bind", "0.0.0.0:8000", "--workers", "${WEB_CONCURRENCY:-4}"]Under high load, monitor CPU usage per worker (
docker stats) and pending request count in Gunicorn logs (--access-logfile -) before increasing the worker count.Orchestrate Services with Docker Compose
In
docker-compose.yml, declare theapiservice, a PostgreSQLdbservice, and aredisservice. Usedepends_onwithcondition: service_healthyso the API waits for the database to be ready before starting:services: api: build: . env_file: .env ports: ["8000:8000"] depends_on: db: condition: service_healthy redis: condition: service_started healthcheck: test: ["CMD", "curl", "-f", "http://localhost:8000/health"] interval: 30s retries: 3 db: image: postgres:16-alpine environment: POSTGRES_USER: user POSTGRES_PASSWORD: password POSTGRES_DB: appdb volumes: ["pgdata:/var/lib/postgresql/data"] healthcheck: test: ["CMD", "pg_isready", "-U", "user"] interval: 10s retries: 5 redis: image: redis:7-alpine volumes: pgdata:Place Nginx as a Reverse Proxy
Configure a server block with
proxy_pass http://api:8000and, for WebSockets, add the necessary headers. Disable buffering for streaming if needed:server { listen 443 ssl; server_name api.yourdomain.com; location / { proxy_pass http://api:8000; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; proxy_buffering off; } }Secure with a TLS Certificate
Obtain a Let's Encrypt certificate via Certbot or a companion container, force HTTPS, and redirect HTTP. Remember to restrict access to
/docsand/openapi.jsonin production if the API is not public.Launch and Test
Start with
docker compose up -d, check health via the/healthendpoint you added, and test the interactive documentation. Monitor the Uvicorn logs to adjust the number of workers under real load.
For long-running tasks (sending emails, image processing, slow external calls), do not block the event loop: offload them to a background worker with BackgroundTasks for simple cases, or to Celery/ARQ with Redis for heavy processing. An API that responds quickly and delegates asynchronous work stays responsive even under high concurrency.
Database and Migrations
For FastAPI applications in production, asyncpg combined with SQLAlchemy 2.x (async mode) offers an excellent balance of performance and maintainability. Configure the connection pool according to expected load:
from sqlalchemy.ext.asyncio import create_async_engine, AsyncSession
from sqlalchemy.orm import sessionmaker
engine = create_async_engine(
DATABASE_URL,
pool_size=10, # persistent connections
max_overflow=20, # additional connections under peak
pool_timeout=30, # seconds before error
pool_recycle=1800, # recycle after 30 min to avoid PostgreSQL timeouts
echo=False,
)
AsyncSessionLocal = sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)Inject the session into your routes via a FastAPI dependency:
async def get_db():
async with AsyncSessionLocal() as session:
yield sessionFor schema migrations, Alembic is the standard tool. Initialize it in your project (alembic init migrations) then configure alembic.ini to point to your DATABASE_URL. Essential commands:
# Generate a migration from your SQLAlchemy models
alembic revision --autogenerate -m "add_users_table"
# Apply pending migrations
alembic upgrade head
# Check current state
alembic currentIn production, run alembic upgrade head before starting the API — add it as an init command in your docker-compose.yml or in a separate startup script. Never run migrations from multiple instances simultaneously.
CI/CD and Automated Deployments
A simple GitHub Actions pipeline is sufficient for most FastAPI projects: build the image, push to a registry, and deploy to the VPS via SSH with zero-downtime.
# .github/workflows/deploy.yml
name: Deploy FastAPI
on:
push:
branches: [main]
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build & push Docker image
run: |
docker build -t ghcr.io/${{ github.repository }}:${{ github.sha }} .
echo ${{ secrets.GITHUB_TOKEN }} | docker login ghcr.io -u ${{ github.actor }} --password-stdin
docker push ghcr.io/${{ github.repository }}:${{ github.sha }}
- name: Deploy via SSH
uses: appleboy/ssh-action@v1
with:
host: ${{ secrets.VPS_HOST }}
username: deploy
key: ${{ secrets.VPS_SSH_KEY }}
script: |
cd /opt/myapp
export IMAGE_TAG=${{ github.sha }}
docker compose pull api
docker compose up -d --no-deps api
docker compose exec api alembic upgrade headThe --no-deps flag on docker compose up restarts only the api container without touching db or redis, ensuring zero-downtime for third-party services. For more robust deployments with instant rollback, tag your images with the commit SHA and keep the last 3 versions locally.
Monitoring and Logs
Access Uvicorn logs directly via Docker:
# Real-time logs
docker compose logs -f api
# Last 100 lines
docker compose logs --tail=100 apiFor structured JSON logs (indexable by an aggregator like Loki or Elasticsearch), replace the standard Python logger with loguru or structlog:
from loguru import logger
import sys
logger.remove()
logger.add(sys.stdout, format="{time} {level} {message}", serialize=True)
@app.middleware("http")
async def log_requests(request, call_next):
logger.info("request", method=request.method, url=str(request.url))
response = await call_next(request)
logger.info("response", status=response.status_code)
return responseTo expose Prometheus metrics, add a /metrics endpoint with prometheus-fastapi-instrumentator:
from prometheus_fastapi_instrumentator import Instrumentator
Instrumentator().instrument(app).expose(app, endpoint="/metrics")Then restrict /metrics to your internal network from Nginx (allow 127.0.0.1; deny all;) to avoid publicly exposing metrics.
Troubleshooting Common Errors
Worker timed out (pid XXXXX) — Gunicorn kills a worker that did not respond within the --timeout delay. Common causes: a synchronous route (def instead of async def) blocks the event loop, or a blocking operation (file read, network call) runs without asyncio.run_in_executor. Increasing --timeout masks the problem; the correct fix is to make the route async or offload blocking work.
Address already in use (port 8000) — A previous Gunicorn process is still active. In Docker: docker compose down then docker compose up -d forces container recreation and releases the port. Check with docker ps -a that no zombie container is running.
502 Bad Gateway after a deployment — Nginx forwarded the request to the API during its restart. Wait 10–15 seconds for the Docker healthcheck to confirm the new container is healthy before considering the deployment complete. Add proxy_read_timeout 60; in Nginx to absorb slow startups.
CORS in production: No 'Access-Control-Allow-Origin' header — FastAPI serves the origins declared in CORSMiddleware. Verify that CORS_ORIGINS in your .env contains exactly the frontend URL (protocol + domain + port if non-standard, no trailing slash). A value of ["*"] works in development but must never go to production.
DB connection refused at startup — The API starts before PostgreSQL is ready. depends_on with condition: service_healthy in Docker Compose solves this. In Kubernetes or deployments without Compose, implement a retry loop with exponential backoff in the application startup code.
Deploy FastAPI in One Click from the ServOrbit Marketplace
The ServOrbit Marketplace offers a preconfigured FastAPI Stack template: Python 3.12, Uvicorn, Gunicorn, Nginx, and PostgreSQL 16 are installed automatically on your VPS from the order form. No SSH session for the base setup — you arrive directly at the repository clone step and start Uvicorn. The marketplace page contains the exact specifications, use cases, and a first-login guide.