Deployment guide

Install Apache Airflow on VPS: Docker Compose and PostgreSQL

Deploy on a VPS Cloud →

Tutorial

Install Apache Airflow on VPS: Docker Compose and PostgreSQL

Automation12 min read9 steps

Apache Airflow is the open-source standard for orchestrating data pipelines, ETL workflows and ML tasks in Python. Self-hosted on a VPS, it frees you from the cost of a managed service (Cloud Composer, MWAA) while keeping full control over your data and dependencies. This guide covers everything from zero to a production-ready setup: PostgreSQL as metadata backend (never SQLite on Docker volumes), CeleryExecutor for task distribution, encrypted secrets and connections, and the right reflexes for fast debugging.

Contents· Why host Apache Airflow on your own VPS1/13
  1. 01Why host Apache Airflow on your own VPS
  2. 02What Airflow orchestrates in practice
  3. 03VPS requirements and technical stack
  4. 04Airflow Docker Compose architecture
  5. 05Deploy Airflow 3.x with Docker Compose
  6. 06Configure PostgreSQL as the metadata backend
  7. 07Switch to CeleryExecutor for multiple workers
  8. 08Manage connections and secret variables
  9. 09Create and test your first Python DAG
  10. 10Secure your Airflow installation
  11. 11LocalExecutor vs CeleryExecutor vs KubernetesExecutor
  12. 12Update Airflow without downtime
  13. 13Debug and read Airflow logs

Why host Apache Airflow on your own VPS

Cloud orchestrators like Cloud Composer (Google) or MWAA (Amazon) bundle Airflow as a managed service, but they bill the environment by the hour — regardless of your actual workload. Self-hosting on a VPS means you only pay for the server, and your DAGs run close to your internal data sources without external network transit. The freedom goes further: you choose the Airflow version, Python providers, database connections, and executor (Local, Celery, Kubernetes). For pipelines processing sensitive data or regulated environments, staying on-premise with a dedicated VPS is not an option but a requirement. Cost-wise, a 4 vCPU / 8 GB RAM VPS is a fraction of a managed environment for the same daily workload.

What Airflow orchestrates in practice

  • Data pipelines and ETL — extract from APIs or databases, transform and load into a data warehouse.
  • ML workflows — model training, evaluation, automated deployment with step dependencies.
  • Complex scheduled tasks — backfills, catchups, per-task configurable retry on failure.
  • Third-party integrations — 80+ official providers: PostgreSQL, MySQL, S3, BigQuery, dbt, Spark, Kubernetes, HTTP.
  • Reporting pipelines — automated periodic report generation and delivery via email or Slack.
  • Microservice orchestration — trigger remote jobs via API and wait for their result before proceeding.

VPS requirements and technical stack

Airflow is the most resource-intensive stack in this category in CeleryExecutor mode, as it runs multiple services in parallel. For production: minimum 4 vCPU and 8 GB RAM — start with a Power VPS or a dedicated VPS. For development or testing: 2 vCPU and 4 GB RAM are enough in LocalExecutor mode (no Celery or Redis). Plan for 40 GB SSD for logs and metadata. Software requirements: Docker 24+ and Docker Compose v2 (check with docker compose version, not legacy docker-compose v1), Ubuntu 22.04 LTS or Debian 12, and a domain name pointing to your VPS for the HTTPS reverse proxy.

Airflow Docker Compose architecture

Airflow's official docker-compose.yaml deploys six services forming a complete architecture. The webserver serves the graphical interface on port 8080: DAG visualization, manual triggering, log inspection. The scheduler is the system's core: it parses DAGs, schedules tasks according to their schedule and dependencies, and queues them for workers. The worker executes tasks assigned by the scheduler — you can run multiple workers for parallelization. The triggerer handles deferrable tasks (async sensors since Airflow 2.2+) without blocking worker slots. PostgreSQL is Airflow's metadata database: DAG run state, structured logs, XComs, variables and connections — it is the most critical component. Redis serves as the message broker between the scheduler and Celery workers: each ready task transits through a Redis queue before being picked up by an available worker.

Deploy Airflow 3.x with Docker Compose

  1. Fetch the official docker-compose and create directories

    Create the working directory, download the reference file and prepare the volume-mounted folders:

    mkdir -p /opt/airflow && cd /opt/airflow
    curl -LfO 'https://airflow.apache.org/docs/apache-airflow/stable/docker-compose.yaml'
    mkdir -p ./dags ./logs ./plugins ./config

    These directories must exist before starting the containers. If Docker creates them, it assigns root ownership and Airflow cannot write to them.

  2. Configure the .env file (UID, PostgreSQL, secret key)

    Create /opt/airflow/.env with the essential variables:

    echo "AIRFLOW_UID=$(id -u)" > .env
    echo "AIRFLOW__DATABASE__SQL_ALCHEMY_CONN=postgresql+psycopg2://airflow:airflow@postgres/airflow" >> .env
    echo "AIRFLOW__CORE__FERNET_KEY=$(python3 -c 'from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())')" >> .env

    AIRFLOW_UID aligns volume permissions with your system user. AIRFLOW__DATABASE__SQL_ALCHEMY_CONN forces PostgreSQL — making it explicit in .env avoids any ambiguity. FERNET_KEY encrypts connections and variables stored in the database — generate it once and back it up: if you lose it, all your secret connections become unreadable.

  3. Initialize the database

    Apply migrations and create the first admin account:

    docker compose up airflow-init

    The airflow-init container connects to PostgreSQL, applies all schema migrations and creates the airflow user (password airflow). Wait for exit code 0 — any other code indicates a PostgreSQL connection issue (not yet healthy) or a missing image. Change the password on first login via the UI or docker compose exec airflow-webserver airflow users reset-password.

  4. Start all services

    Launch the full stack in the background:

    docker compose up -d
    docker compose ps

    All containers should reach healthy state within one to three minutes. If a service stays in starting, check its logs:

    docker compose logs airflow-scheduler --tail=50

    The webserver is accessible at http://localhost:8080. Do not expose this port publicly — use a reverse proxy.

  5. Create a custom admin account

    The default airflow/airflow account created by airflow-init is for first-boot only. Create your permanent account:

    docker compose exec airflow-webserver airflow users create \
      --username admin \
      --firstname Firstname \
      --lastname Lastname \
      --role Admin \
      --email [email protected] \
      --password YourStrongPassword

    Then delete the default account via Admin → Users in the web interface.

  6. Expose Airflow behind an HTTPS reverse proxy

    Never expose port 8080 directly — the Airflow UI displays connections, variables and sensitive logs. With Caddy (recommended for automatic TLS):

    apt install -y caddy
    cat > /etc/caddy/Caddyfile << 'EOF'
    airflow.yourdomain.com {
        reverse_proxy localhost:8080
    }
    EOF
    systemctl reload caddy
    ufw deny 8080

    Caddy automatically obtains and renews the Let's Encrypt certificate. With nginx, create a standard vhost with proxy_pass http://127.0.0.1:8080 and block the port with ufw deny 8080.

Configure PostgreSQL as the metadata backend

Never use SQLite in production with Airflow on Docker. SQLite is not designed for concurrent multi-process access: the scheduler, webserver, and workers access the database simultaneously, and SQLite on Docker volumes has locking limitations that cause corruption and data loss — the Airflow project explicitly documents this warning since version 2.0. The official docker-compose.yaml already includes a PostgreSQL 13 service, but verify your .env explicitly forces the connection:

AIRFLOW__DATABASE__SQL_ALCHEMY_CONN=postgresql+psycopg2://airflow:airflow@postgres/airflow

For production, customize the PostgreSQL password. In docker-compose.yaml, find the postgres section and modify POSTGRES_PASSWORD. Update AIRFLOW__DATABASE__SQL_ALCHEMY_CONN accordingly. Back up the PostgreSQL database regularly: docker compose exec postgres pg_dump -U airflow airflow > backup_$(date +%Y%m%d).sql. If you already manage an external PostgreSQL server, remove the postgres service from the compose and point SQL_ALCHEMY_CONN to your instance — this is the recommended configuration for multi-VPS setups.

Switch to CeleryExecutor for multiple workers

The LocalExecutor runs tasks as subprocess of the scheduler — simple, but limited to a single node. The CeleryExecutor decouples execution: the scheduler pushes tasks into a Redis queue, and one or more workers consume them independently. The official docker-compose.yaml enables CeleryExecutor by default. To scale out, simply replicate the worker service:

docker compose up -d --scale airflow-worker=3

This starts three workers consuming the same Redis queue in parallel. Each worker handles parallelism simultaneous tasks (configurable via AIRFLOW__CORE__PARALLELISM, default 32). To distribute workers across multiple VPS, point them all to the same Redis and PostgreSQL — that is standard Celery architecture. Add AIRFLOW__CELERY__WORKER_CONCURRENCY=8 in .env to fine-tune per-worker concurrency based on available resources on each node.

Manage connections and secret variables

Airflow encrypts connections and sensitive variables in the database using the FERNET_KEY configured in .env. Via the UI: Admin → Connections → Add Connection — fill in the full URI or individual fields (host, login, password, port). Values are Fernet-encrypted before being written to the database. Via environment variables (recommended for CI/CD): an env variable prefixed with AIRFLOW_CONN_ overrides any connection stored in the database. Example in .env:

AIRFLOW_CONN_MY_POSTGRES=postgresql://user:password@host:5432/dbname

The connection name in your DAGs is then my_postgres (prefix and uppercase are stripped). Secret Backend: for Kubernetes or cloud environments, configure AIRFLOW__SECRETS__BACKEND=airflow.providers.amazon.aws.secrets.secrets_manager.SecretsManagerBackend (or HashiCorp Vault, GCP Secret Manager) — Airflow will resolve connections from your vault rather than from PostgreSQL. Variables: Admin → Variables in the UI, or via env AIRFLOW_VAR_VARIABLE_NAME. Avoid storing secrets in unencrypted Variables — prefer Connections or Secret Backend.

Create and test your first Python DAG

  1. Create the DAG file in the /dags folder

    Create /opt/airflow/dags/my_first_dag.py:

    from airflow.sdk import DAG, task
    from datetime import datetime
    
    with DAG(
        dag_id='my_first_dag',
        schedule='@daily',
        start_date=datetime(2025, 1, 1),
        catchup=False,
        tags=['example'],
    ) as dag:
    
        @task
        def extract():
            return {'rows': 42}
    
        @task
        def transform(data: dict):
            print(f"Processing {data['rows']} rows")
            return data['rows'] * 2
    
        @task
        def load(result: int):
            print(f"Final result: {result}")
    
        load(transform(extract()))

    Note the use of schedule (not schedule_interval, removed in Airflow 3).

  2. Verify the scheduler parses the DAG without errors

    The scheduler automatically detects and parses .py files in the dags folder. Check manually:

    docker compose exec airflow-scheduler airflow dags list | grep my_first_dag

    If the DAG does not appear after 30 seconds, look for parsing errors:

    docker compose logs airflow-scheduler | grep -i 'error\|exception' | tail -20

    A Python import error in the DAG file blocks all detection — fix the syntax and the scheduler will automatically retry parsing.

  3. Trigger a manual run and inspect logs

    In the Airflow UI (https://airflow.yourdomain.com), navigate to DAGs, find my_first_dag, enable it (toggle), then click ▶ Trigger DAG. Follow the execution in real time in the Graph or Grid view. Click on a task, then Logs to see the function's standard output. From the terminal:

    docker compose exec airflow-scheduler airflow dags trigger my_first_dag
    docker compose exec airflow-scheduler airflow dags state my_first_dag $(date +%Y-%m-%dT%H:%M:%S+00:00)

Secure your Airflow installation

An Airflow installation potentially exposes your database connections, variables and execution logs. Apply these measures from the start. Authentication: Airflow 3 uses FAB (Flask-AppBuilder) with password authentication by default. For an organization, connect Airflow to your directory via LDAP, or enable OAuth2 (GitHub, Google, Azure AD) by configuring the corresponding FAB provider. Disable example DAGs: add to .env:

AIRFLOW__CORE__LOAD_EXAMPLES=False

These dummy DAGs clutter the UI and can be misleading. Firewall on port 8080: ufw deny 8080 — only the reverse proxy should access it from localhost. No Fernet key in logs: never log os.environ, which would print the FERNET_KEY in plain text. Secret rotation: if you need to change the FERNET_KEY, use airflow db rotate-fernet-key to re-encrypt all stored connections before updating the variable.

LocalExecutor vs CeleryExecutor vs KubernetesExecutor

Scroll the table

CriterionLocalExecutorCeleryExecutorKubernetesExecutor
Required resources2 vCPU / 4 GB RAM4 vCPU / 8 GB RAM + RedisKubernetes cluster required
ParallelismLimited to scheduler CPU countHorizontal — as many workers as neededOne pod per task — near-unbounded parallelism
Task isolationShared subprocessesSeparate workers, same imageIsolated pod per task, configurable image
Operational complexitySimple — no RedisModerate — Redis to maintainHigh — Kubernetes required
Ideal use caseDev, tests, light pipelinesVPS production, variable loadCloud-native, strong isolation required

Update Airflow without downtime

For a rolling update without interrupting running DAGs: start by updating the image in docker-compose.yaml (e.g. apache/airflow:3.3.1). Stop only the webserver and scheduler, apply migrations, then restart service by service. Celery workers continue executing already-started tasks during the migration:

docker compose stop airflow-webserver airflow-scheduler
docker compose pull
docker compose up -d airflow-webserver airflow-scheduler
docker compose exec airflow-scheduler airflow db upgrade
docker compose up -d

Always check the official CHANGELOG between minor versions — some schema migrations are irreversible. Keep a fresh PostgreSQL backup before any version upgrade.

Debug and read Airflow logs

Airflow generates logs at multiple levels and understanding their organization greatly speeds up debugging. Task logs: each task writes to ./logs/dag_id/run_id/task_id/. Accessible from the UI (click on task → Logs) or from the command line:

docker compose exec airflow-scheduler airflow tasks logs my_first_dag extract 2025-01-01

Scheduler logs: the most important for diagnosing DAGs that do not trigger:

docker compose logs airflow-scheduler --follow --tail=100

Remote logging: for multi-worker setups or long-retention logs, offload to S3 or GCS:

AIRFLOW__LOGGING__REMOTE_LOGGING=True
AIRFLOW__LOGGING__REMOTE_BASE_LOG_FOLDER=s3://my-bucket/airflow-logs
AIRFLOW__LOGGING__REMOTE_LOG_CONN_ID=my_aws

Silent errors: if a DAG disappears from the UI without a message, look for scheduler import errors:

docker compose exec airflow-webserver airflow dags list-import-errors

Log purge: set AIRFLOW__LOG_RETENTION_DAYS=30 to avoid disk saturation — the ./logs folder can quickly reach several gigabytes on frequent pipelines.

The ideal Cloud VPS for Apache Airflow

Airflow 3 in Celery mode demands RAM and several containers. The ServOrbit Cloud VPS provides the resources, preconfigured Docker and automatic SSL to host your data scheduler self-hosted, at a controlled cost.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab