Deployment guide

Hosting an open source AI application on your VPS

Deploy on a VPS Cloud →

Tutorial

Hosting an open source AI application on your VPS

Artificial Intelligence10 min read5 steps

The open source AI tooling ecosystem has exploded over the past two years. Open WebUI, Dify, Flowise, AnythingLLM, n8n with AI nodes, OpenClaw… Each tool meets a specific need. But they all share one common requirement: they run better, more securely, and at a lower cost on your own infrastructure.

Contents· Why a VPS beats AWS or GCP for your AI tools1/11
  1. 01Why a VPS beats AWS or GCP for your AI tools
  2. 026 concrete advantages of a VPS for hosting your AI tools
  3. 03Choosing your AI tool by use case — 2026 updated guide
  4. 04Expanded selection guide — 2026 state
  5. 05VPS requirements per AI tool — RAM, CPU and GPU
  6. 06Choosing the Ollama quantization
  7. 07Deploy a multi-tool AI stack with Docker Compose
  8. 08Configuring reverse proxies to access each tool
  9. 09Securing access to your AI tools
  10. 10Troubleshooting — common errors
  11. 11Going further — related articles

Why a VPS beats AWS or GCP for your AI tools

Managed cloud services (AWS, GCP, Azure) are powerful, but their pay-per-resource billing model becomes expensive quickly once a container runs continuously. A simple Open WebUI Docker container running 24/7 on AWS EC2 (t3.large, 2 vCPU / 8 GB) costs roughly 60–80 € per month before storage and data transfer. A VPS with the same resources costs a fraction of that, with a fixed and predictable monthly price.

Three other advantages matter more than cost for AI tools:

Local API latency. When Open WebUI calls Ollama, both services are on the same machine: latency is in the millisecond range. On a distributed cloud architecture, each exchange crosses the internal network — measurable friction for token-by-token generation.

Data sovereignty. Your prompts, RAG documents and conversation logs never leave your infrastructure. Non-negotiable for professional use cases involving sensitive data.

Environment control. Full root access: you choose the GPU driver, CUDA versions, and network configuration. Managed cloud services abstract this layer — convenient to start, limiting for AI stacks that require fine-tuned adjustments.

6 concrete advantages of a VPS for hosting your AI tools

  • Fixed and predictable cost — no surprise billing based on model activity
  • Data that stays on your infrastructure — no transfer to third parties
  • Full root access — configure the environment exactly as you want
  • Docker internal network — Open WebUI and Ollama on the same machine share the Docker bridge network, without crossing the public network
  • Ability to run multiple tools on a single server with Docker Compose
  • No restrictions on number of requests, users or generated tokens

Choosing your AI tool by use case — 2026 updated guide

The choice depends on what you want to accomplish. Open WebUI (v0.11.x) has established itself as the reference chat interface for Ollama and OpenAI-compatible APIs: it supports tools, documents, pipelines and user management. Dify (v1.15.x) stands out for building LLM applications with a complete visual interface — workflow, RAG, agents and integrated monitoring. n8n (v1.123.x) remains the reference for hybrid workflow automation combining AI nodes and classic business integrations. LangFlow (v1.11.x) and Flowise (v3.1.x) target the same niche — visual LLM pipelines based on LangChain — with different philosophies: LangFlow focuses on Python customization, Flowise on simplicity. LangFuse (v3.x) is not a generation tool but an observability tool: it traces LLM calls, measures latency and response quality, and integrates in a few lines with the other tools on this list.

Expanded selection guide — 2026 state

Scroll the table

NeedRecommended toolMin. VPS resources
Chat interface with an LLM (Ollama or OpenAI API)Open WebUI v0.11.x2 GB RAM, 2 vCPU
Build LLM applications visually (RAG, agents)Dify v1.15.x4 GB RAM, 2 vCPU
Automate workflows mixing AI and business toolsn8n v1.123.x + AI nodes2 GB RAM, 2 vCPU
RAG pipelines and no-code chatbots with LangChainFlowise v3.1.x2 GB RAM, 1 vCPU
Python-customizable LLM pipelinesLangFlow v1.11.x3 GB RAM, 2 vCPU
LLM call observability and traceabilityLangFuse v3.x2 GB RAM, 1 vCPU
Personal knowledge base (PKM) with AIKarakeep or Anytype self-hosted2 GB RAM, 1 vCPU
Local chat interface without web UI (CLI)Jan.ai (server mode)2 GB RAM, 2 vCPU
Multiple tools on a single VPSCoolify or Dokploy8 GB RAM, 4 vCPU

VPS requirements per AI tool — RAM, CPU and GPU

Needs vary significantly from one tool to another. Here are the resources to plan for by use case:

Interfaces and orchestrators (without local model). Open WebUI alone uses less than 500 MB of RAM: it is stateless and delegates inference to Ollama or an external API. Dify is more resource-intensive — its Docker Compose stack includes multiple services (worker, API, web, vector database, Redis, PostgreSQL) and uses approximately 3.5–4 GB at idle. n8n and Flowise are lightweight: plan 256–512 MB per service.

Ollama with local models. This is where resources really matter. The baseline rule applies to default Q4 quantization (what Ollama downloads if you don't specify): a 7B model requires 4–5 GB RAM (8 GB recommended to leave headroom for the OS), a 13B around 8–9 GB (16 GB recommended), and a 70B between 38 and 48 GB depending on context length. Without a dedicated GPU, inference runs on CPU — functional for testing, but generating 100 tokens can take 30 to 60 seconds on a 7B CPU-only.

Dedicated GPU. For daily use with 13B+ models, or for handling multiple simultaneous requests, a GPU server is necessary. VPS instances with NVIDIA GPU allow storing the model entirely in VRAM and reduce generation to 1–3 seconds per 100 tokens.

Choosing the Ollama quantization

Ollama downloads the Q4_K_M variant of a model by default — a quality/memory tradeoff suited to most uses. If your VPS has exactly 8 GB RAM for a 7B, prefer Q4_K_S (slightly less precise, ~200 MB less) to leave headroom. The command ollama run llama3.2:7b-instruct-q4_K_S explicitly selects this variant.

Deploy a multi-tool AI stack with Docker Compose

  1. Prepare the VPS

    Connect via SSH and install Docker Engine and the Compose plugin:

    curl -fsSL https://get.docker.com | sh
    systemctl enable --now docker

    Verify Docker works: docker run --rm hello-world.

  2. Create the project structure

    Create a folder for your AI stack:

    mkdir -p /opt/ai-stack && cd /opt/ai-stack

    This folder will contain your docker-compose.yml and data volumes.

  3. Write the docker-compose.yml file

    Here is an example stack combining Open WebUI, Ollama and n8n:

    services:
      ollama:
        image: ollama/ollama:latest
        volumes:
          - ollama_data:/root/.ollama
        restart: unless-stopped
    
      open-webui:
        image: ghcr.io/open-webui/open-webui:main
        ports:
          - "3000:8080"
        environment:
          - OLLAMA_BASE_URL=http://ollama:11434
        volumes:
          - open_webui_data:/app/backend/data
        depends_on:
          - ollama
        restart: unless-stopped
    
      n8n:
        image: n8nio/n8n:latest
        ports:
          - "5678:5678"
        environment:
          - N8N_HOST=n8n.yourdomain.com
          - N8N_PROTOCOL=https
          - WEBHOOK_URL=https://n8n.yourdomain.com/
        volumes:
          - n8n_data:/home/node/.n8n
        restart: unless-stopped
    
    volumes:
      ollama_data:
      open_webui_data:
      n8n_data:

    Note that services on the same Docker network communicate using their service name: ollama:11434 is accessible from open-webui without exposing the port on the host.

  4. Start the stack

    Launch all services:

    docker compose up -d

    Check that containers are starting:

    docker compose ps

    Then download a first Ollama model:

    docker compose exec ollama ollama pull llama3.2:latest
  5. Configure the Nginx reverse proxy

    Install Nginx on the host and create a vhost for each tool. Example for Open WebUI:

    server {
        server_name chat.yourdomain.com;
        location / {
            proxy_pass http://127.0.0.1:3000;
            proxy_set_header Host $host;
            proxy_set_header X-Real-IP $remote_addr;
            proxy_set_header Upgrade $http_upgrade;
            proxy_set_header Connection "upgrade";
        }
    }

    Enable HTTPS with Certbot: certbot --nginx -d chat.yourdomain.com.

Configuring reverse proxies to access each tool

With multiple tools on the same VPS, the recommended practice is to assign a subdomain to each service and let a reverse proxy (Nginx or Caddy) route HTTPS traffic to the right port.

Recommended structure:
- chat.yourdomain.com → Open WebUI (port 3000)
- n8n.yourdomain.com → n8n (port 5678)
- dify.yourdomain.com → Dify (port 80 of the web container)
- langfuse.yourdomain.com → LangFuse (port 3000)

Never expose Ollama's port 11434 on the public interface — it has no native authentication. Keep it accessible only from the Docker internal network (OLLAMA_HOST=127.0.0.1 or via Docker bridge network).

Caddy simplifies TLS configuration: it automatically obtains and renews Let's Encrypt certificates. A minimal Caddyfile:

chat.yourdomain.com {
    reverse_proxy localhost:3000
}
n8n.yourdomain.com {
    reverse_proxy localhost:5678
}

caddy run --config /etc/caddy/Caddyfile — and certificates are managed without intervention.

Securing access to your AI tools

Before exposing an AI tool on a public subdomain, check three things:

1. Authentication enabled. Open WebUI creates an admin account on first access — do it immediately. Dify and n8n also require initial setup: do not leave these interfaces open without a password.
2. Port 11434 (Ollama) not exposed. This port has no auth — do not open a firewall rule on it. Check with ss -tlnp | grep 11434: the bind should be on 127.0.0.1 or a Docker interface, not 0.0.0.0.
3. HTTPS mandatory. A valid TLS certificate prevents interception of prompts and authentication tokens in transit. Certbot and Caddy automate renewal.

Troubleshooting — common errors

Ollama: out of memory. The model exceeds available RAM. Check free memory with free -h, then choose a lighter quantization (q4_K_S or q3_K_M) or a smaller model. If the error occurs on a VPS with GPU, VRAM is saturated — check with nvidia-smi.

Open WebUI cannot find Ollama. The configured URL is incorrect or the Docker network does not allow communication. If Open WebUI and Ollama are in the same docker-compose.yml, the correct URL is http://ollama:11434 (service name, not localhost or the host IP). If Ollama is on the host (outside Docker), use http://host.docker.internal:11434 (Linux) or the Docker bridge IP (172.17.0.1 by default).

Port 11434 not accessible from another container. Ollama may be configured to listen only on 127.0.0.1. Add OLLAMA_HOST=0.0.0.0 to the Ollama service environment variables in your Compose file — this exposes it on all container interfaces only, not the public host.

n8n: incorrect webhook URL. n8n builds its webhook URLs from N8N_HOST and WEBHOOK_URL. If these variables don't match your HTTPS subdomain, incoming webhooks will fail. Verify and restart the container after modification: docker compose restart n8n.

Dify: services not starting. Dify uses several interdependent services (PostgreSQL, Redis, worker service). Inspect logs: docker compose logs --tail=50 worker api. The issue is often a pending database migration or an improperly initialized volume.

Going further — related articles

This guide covers setting up a general AI stack. To go deeper on each tool:

- Open WebUI: complete deployment with user management, pipelines and API integrations — see Hosting Open WebUI on a VPS.
- Dify: Docker Compose installation, offline plugin configuration and reverse proxy — see Deploying Dify on a VPS.
- LangFlow: advanced deployment with custom Python backend — see the dedicated article.
- LangFuse: self-hosted LLM observability — see the dedicated LangFuse article.
- Coolify / Dokploy: if you prefer a graphical interface to manage your containers rather than editing Compose files — see Coolify vs Dokploy.

Your AI lab ready in just a few minutes.

Order a ServOrbit Cloud VPS and choose your template from our AI catalog. Automatic installation, root access, dedicated IPv4.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab