Why a VPS beats AWS or GCP for your AI tools
Managed cloud services (AWS, GCP, Azure) are powerful, but their pay-per-resource billing model becomes expensive quickly once a container runs continuously. A simple Open WebUI Docker container running 24/7 on AWS EC2 (t3.large, 2 vCPU / 8 GB) costs roughly 60–80 € per month before storage and data transfer. A VPS with the same resources costs a fraction of that, with a fixed and predictable monthly price.
Three other advantages matter more than cost for AI tools:
Local API latency. When Open WebUI calls Ollama, both services are on the same machine: latency is in the millisecond range. On a distributed cloud architecture, each exchange crosses the internal network — measurable friction for token-by-token generation.
Data sovereignty. Your prompts, RAG documents and conversation logs never leave your infrastructure. Non-negotiable for professional use cases involving sensitive data.
Environment control. Full root access: you choose the GPU driver, CUDA versions, and network configuration. Managed cloud services abstract this layer — convenient to start, limiting for AI stacks that require fine-tuned adjustments.
6 concrete advantages of a VPS for hosting your AI tools
- Fixed and predictable cost — no surprise billing based on model activity
- Data that stays on your infrastructure — no transfer to third parties
- Full root access — configure the environment exactly as you want
- Docker internal network — Open WebUI and Ollama on the same machine share the Docker bridge network, without crossing the public network
- Ability to run multiple tools on a single server with Docker Compose
- No restrictions on number of requests, users or generated tokens
Choosing your AI tool by use case — 2026 updated guide
The choice depends on what you want to accomplish. Open WebUI (v0.11.x) has established itself as the reference chat interface for Ollama and OpenAI-compatible APIs: it supports tools, documents, pipelines and user management. Dify (v1.15.x) stands out for building LLM applications with a complete visual interface — workflow, RAG, agents and integrated monitoring. n8n (v1.123.x) remains the reference for hybrid workflow automation combining AI nodes and classic business integrations. LangFlow (v1.11.x) and Flowise (v3.1.x) target the same niche — visual LLM pipelines based on LangChain — with different philosophies: LangFlow focuses on Python customization, Flowise on simplicity. LangFuse (v3.x) is not a generation tool but an observability tool: it traces LLM calls, measures latency and response quality, and integrates in a few lines with the other tools on this list.
Expanded selection guide — 2026 state
Scroll the table
| Need | Recommended tool | Min. VPS resources |
|---|---|---|
| Chat interface with an LLM (Ollama or OpenAI API) | Open WebUI v0.11.x | 2 GB RAM, 2 vCPU |
| Build LLM applications visually (RAG, agents) | Dify v1.15.x | 4 GB RAM, 2 vCPU |
| Automate workflows mixing AI and business tools | n8n v1.123.x + AI nodes | 2 GB RAM, 2 vCPU |
| RAG pipelines and no-code chatbots with LangChain | Flowise v3.1.x | 2 GB RAM, 1 vCPU |
| Python-customizable LLM pipelines | LangFlow v1.11.x | 3 GB RAM, 2 vCPU |
| LLM call observability and traceability | LangFuse v3.x | 2 GB RAM, 1 vCPU |
| Personal knowledge base (PKM) with AI | Karakeep or Anytype self-hosted | 2 GB RAM, 1 vCPU |
| Local chat interface without web UI (CLI) | Jan.ai (server mode) | 2 GB RAM, 2 vCPU |
| Multiple tools on a single VPS | Coolify or Dokploy | 8 GB RAM, 4 vCPU |
VPS requirements per AI tool — RAM, CPU and GPU
Needs vary significantly from one tool to another. Here are the resources to plan for by use case:
Interfaces and orchestrators (without local model). Open WebUI alone uses less than 500 MB of RAM: it is stateless and delegates inference to Ollama or an external API. Dify is more resource-intensive — its Docker Compose stack includes multiple services (worker, API, web, vector database, Redis, PostgreSQL) and uses approximately 3.5–4 GB at idle. n8n and Flowise are lightweight: plan 256–512 MB per service.
Ollama with local models. This is where resources really matter. The baseline rule applies to default Q4 quantization (what Ollama downloads if you don't specify): a 7B model requires 4–5 GB RAM (8 GB recommended to leave headroom for the OS), a 13B around 8–9 GB (16 GB recommended), and a 70B between 38 and 48 GB depending on context length. Without a dedicated GPU, inference runs on CPU — functional for testing, but generating 100 tokens can take 30 to 60 seconds on a 7B CPU-only.
Dedicated GPU. For daily use with 13B+ models, or for handling multiple simultaneous requests, a GPU server is necessary. VPS instances with NVIDIA GPU allow storing the model entirely in VRAM and reduce generation to 1–3 seconds per 100 tokens.
Choosing the Ollama quantization
Ollama downloads the Q4_K_M variant of a model by default — a quality/memory tradeoff suited to most uses. If your VPS has exactly 8 GB RAM for a 7B, prefer Q4_K_S (slightly less precise, ~200 MB less) to leave headroom. The command ollama run llama3.2:7b-instruct-q4_K_S explicitly selects this variant.
Deploy a multi-tool AI stack with Docker Compose
Prepare the VPS
Connect via SSH and install Docker Engine and the Compose plugin:
curl -fsSL https://get.docker.com | sh systemctl enable --now dockerVerify Docker works:
docker run --rm hello-world.Create the project structure
Create a folder for your AI stack:
mkdir -p /opt/ai-stack && cd /opt/ai-stackThis folder will contain your
docker-compose.ymland data volumes.Write the docker-compose.yml file
Here is an example stack combining Open WebUI, Ollama and n8n:
services: ollama: image: ollama/ollama:latest volumes: - ollama_data:/root/.ollama restart: unless-stopped open-webui: image: ghcr.io/open-webui/open-webui:main ports: - "3000:8080" environment: - OLLAMA_BASE_URL=http://ollama:11434 volumes: - open_webui_data:/app/backend/data depends_on: - ollama restart: unless-stopped n8n: image: n8nio/n8n:latest ports: - "5678:5678" environment: - N8N_HOST=n8n.yourdomain.com - N8N_PROTOCOL=https - WEBHOOK_URL=https://n8n.yourdomain.com/ volumes: - n8n_data:/home/node/.n8n restart: unless-stopped volumes: ollama_data: open_webui_data: n8n_data:Note that services on the same Docker network communicate using their service name:
ollama:11434is accessible fromopen-webuiwithout exposing the port on the host.Start the stack
Launch all services:
docker compose up -dCheck that containers are starting:
docker compose psThen download a first Ollama model:
docker compose exec ollama ollama pull llama3.2:latestConfigure the Nginx reverse proxy
Install Nginx on the host and create a vhost for each tool. Example for Open WebUI:
server { server_name chat.yourdomain.com; location / { proxy_pass http://127.0.0.1:3000; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; } }Enable HTTPS with Certbot:
certbot --nginx -d chat.yourdomain.com.
Configuring reverse proxies to access each tool
With multiple tools on the same VPS, the recommended practice is to assign a subdomain to each service and let a reverse proxy (Nginx or Caddy) route HTTPS traffic to the right port.
Recommended structure:
- chat.yourdomain.com → Open WebUI (port 3000)
- n8n.yourdomain.com → n8n (port 5678)
- dify.yourdomain.com → Dify (port 80 of the web container)
- langfuse.yourdomain.com → LangFuse (port 3000)
Never expose Ollama's port 11434 on the public interface — it has no native authentication. Keep it accessible only from the Docker internal network (OLLAMA_HOST=127.0.0.1 or via Docker bridge network).
Caddy simplifies TLS configuration: it automatically obtains and renews Let's Encrypt certificates. A minimal Caddyfile:
chat.yourdomain.com {
reverse_proxy localhost:3000
}
n8n.yourdomain.com {
reverse_proxy localhost:5678
}caddy run --config /etc/caddy/Caddyfile — and certificates are managed without intervention.
Securing access to your AI tools
Before exposing an AI tool on a public subdomain, check three things:
1. Authentication enabled. Open WebUI creates an admin account on first access — do it immediately. Dify and n8n also require initial setup: do not leave these interfaces open without a password.
2. Port 11434 (Ollama) not exposed. This port has no auth — do not open a firewall rule on it. Check with ss -tlnp | grep 11434: the bind should be on 127.0.0.1 or a Docker interface, not 0.0.0.0.
3. HTTPS mandatory. A valid TLS certificate prevents interception of prompts and authentication tokens in transit. Certbot and Caddy automate renewal.
Troubleshooting — common errors
Ollama: out of memory. The model exceeds available RAM. Check free memory with free -h, then choose a lighter quantization (q4_K_S or q3_K_M) or a smaller model. If the error occurs on a VPS with GPU, VRAM is saturated — check with nvidia-smi.
Open WebUI cannot find Ollama. The configured URL is incorrect or the Docker network does not allow communication. If Open WebUI and Ollama are in the same docker-compose.yml, the correct URL is http://ollama:11434 (service name, not localhost or the host IP). If Ollama is on the host (outside Docker), use http://host.docker.internal:11434 (Linux) or the Docker bridge IP (172.17.0.1 by default).
Port 11434 not accessible from another container. Ollama may be configured to listen only on 127.0.0.1. Add OLLAMA_HOST=0.0.0.0 to the Ollama service environment variables in your Compose file — this exposes it on all container interfaces only, not the public host.
n8n: incorrect webhook URL. n8n builds its webhook URLs from N8N_HOST and WEBHOOK_URL. If these variables don't match your HTTPS subdomain, incoming webhooks will fail. Verify and restart the container after modification: docker compose restart n8n.
Dify: services not starting. Dify uses several interdependent services (PostgreSQL, Redis, worker service). Inspect logs: docker compose logs --tail=50 worker api. The issue is often a pending database migration or an improperly initialized volume.
Going further — related articles
This guide covers setting up a general AI stack. To go deeper on each tool:
- Open WebUI: complete deployment with user management, pipelines and API integrations — see Hosting Open WebUI on a VPS.
- Dify: Docker Compose installation, offline plugin configuration and reverse proxy — see Deploying Dify on a VPS.
- LangFlow: advanced deployment with custom Python backend — see the dedicated article.
- LangFuse: self-hosted LLM observability — see the dedicated LangFuse article.
- Coolify / Dokploy: if you prefer a graphical interface to manage your containers rather than editing Compose files — see Coolify vs Dokploy.