Why Self-Host Your LLM Models on a VPS
Self-hosting a model server means keeping your prompts and sensitive data off commercial APIs, removing per-token billing, and setting the model, its version, and its quantization yourself. On a VPS, you expose a private API to your internal applications (chatbots, RAG, code assistants) without any leak to the outside. Ollama stands out for its radical simplicity: one command to download and run a model, a clean API, automatic memory management. LocalAI positions itself as a 'drop-in' replacement for the OpenAI API: it exposes the same endpoints (chat, embeddings, images, audio) and accepts multiple backends and model formats. The choice comes down to a minimalist experience versus the broadest possible compatibility.
The Benefits of a Self-Hosted LLM
- Prompts and confidential data that never leave your VPS
- No per-token billing, predictable cost tied only to the server
- A private API wired directly into your internal apps
- Free choice of the model, its size, and its quantization level
- Compatibility with existing SDKs via an OpenAI-style API
- Ideal for RAG: pair the LLM with your database and a self-hosted search engine
Prerequisites: Quantized Models and Realistic RAM
Without a GPU, stick to lightweight quantized models. A 3B model in GGUF Q4 runs on a VPS with 4 vCPUs and 8 GB of RAM with acceptable latency for testing; a 7B/8B Q4 model needs 8 to 16 GB of RAM and a good CPU to stay usable. Beyond that, on pure CPU, latency becomes prohibitive: for real-time on large models, a VPS with a GPU is essential. Above all, plan for a generous disk: the weights of several models quickly add up to tens of GB. You will need Docker and Compose, a persistent volume for the models, a subdomain if you expose the API, and a reverse proxy with authentication.
Scroll the table
| Criterion | Ollama | LocalAI |
|---|---|---|
| Philosophy | Simplicity, one command per model | Drop-in replacement for the OpenAI API |
| Installation | Very fast, official Docker image ready | More configuration, many backends |
| OpenAI API compatibility | Compatible `/v1/chat/completions` endpoint | Native and very complete, all endpoints |
| Supported models | Official library at ollama.com/library | GGUF, GPTQ, ONNX, exllama2 and more |
| GPU support | NVIDIA, AMD, Apple Metal — auto-detection | NVIDIA, AMD, CPU — broad hardware support |
| Simultaneous multi-model | Yes, automatic load/unload | Yes, isolated backends per model |
| Built-in web interface | No (Open WebUI separate) | No (third-party interface recommended) |
| Ideal use case | Start fast, prototype | Migrate an OpenAI app to self-hosting |
Deploy Ollama (or LocalAI) on a VPS
Provision Storage and the Models Volume
Create a dedicated volume (
/srv/ollama/models) on a sufficiently large disk. Since the weights are large and reusable, they must persist outside the container to avoid re-downloading on every restart.Launch the Container
Start
ollama/ollama(orlocalai/localai), mounting the models volume and binding port11434/8080locally. On a VPS without a GPU, CPU mode is automatic; with a GPU, enable the appropriate runtime.Download a Model
With Ollama, run
docker compose exec ollama ollama pull llama3.2:3b. With LocalAI, declare the model in the gallery or drop the GGUF file into the models folder, then verify its loading in the logs.Test the API
Make a local call:
curl http://127.0.0.1:11434/api/generatefor Ollama, or the OpenAI-compatible endpointPOST /v1/chat/completionsfor LocalAI. Confirm that generation works before any exposure.Expose Behind an Authenticated Reverse Proxy
Proxy llm.yourdomain.com to the local port with Caddy or Nginx for TLS, and add an authentication layer (an API key in a header or basic auth). An LLM API open to the Internet is a costly entry point that should never be left uncontrolled.
Connect Your Applications
Point your existing SDKs to your private
base_url. Since LocalAI exposes the OpenAI API, most libraries work by simply changing the URL and the key; on the Ollama side, use its native API or its compatible endpoint.
Managing Models: ollama pull, list and Deletion
With Ollama, all model management goes through the CLI or the REST API. ollama pull llama3.2 downloads the default variant; ollama pull llama3.1:8b-instruct-q4_K_M pulls a precise quantized variant from the official library at ollama.com/library. ollama list shows local models with their disk size and digest. ollama run mistral opens an interactive session in the terminal. ollama rm llama3.2 deletes a model to free up disk space. With LocalAI, the logic differs: models live in the /models folder mounted inside the container. You download a GGUF (llama.cpp format) or GPTQ file and place it in that directory; LocalAI loads it on startup or on demand. The LocalAI gallery offers ready-made configurations for popular models. A practical point: Ollama automatically unloads unused models to free RAM, whereas LocalAI leaves this to your configuration.
Performance and Memory: Choosing the Right Quantization
Quantization reduces the precision of the weights to shrink model size and memory consumption. Q4_K_M and Q5_K_S are the most common formats in practice. Q4_K_M (4-bit, K method, M size) offers an excellent size/quality trade-off for most workloads: a Llama 3.1 8B model in Q4_K_M uses around 4.7 GB of RAM, versus about 16 GB for the non-quantized FP16 version (source: ollama.com/library/llama3.1). Q5_K_S rises to about 5.5 GB but preserves coherence better on long generations. On an 8 GB VPS, target 3B or 7B models in Q4; on 16 GB, a 7B/8B Q5_K_M stays fluid. On pure CPU, expect 3 to 10 tokens/s depending on the model and core count — acceptable for async use, too slow for interactive streaming. With an NVIDIA GPU, Ollama automatically detects CUDA and offloads layers to the GPU; same logic for AMD (ROCm) and Apple Silicon (Metal). LocalAI provides finer control over GPU layer assignment via backend configuration parameters, useful when GPU and CPU must share the workload.
On a VPS without a GPU, the secret to smoothness is quantization: a model in Q4_K_M offers a very favourable size/quality trade-off and fits in RAM where the non-quantized version collapses. Also limit the context window to the strict minimum (for example 4096 tokens): an oversized context multiplies memory consumption and latency with no real benefit for most tasks.
Integration: Open WebUI, OpenAI API and SDKs
Open WebUI is the reference graphical interface for Ollama: it deploys as a single Docker container and connects to Ollama via the OLLAMA_BASE_URL=http://ollama:11434 environment variable. You get a full chat with history, model selection, conversation management — similar to ChatGPT but hosted on your VPS. For programmatic integrations, Ollama exposes /api/generate (simple generation) and /v1/chat/completions (OpenAI-compatible). An existing Python client reconfigures as: client = openai.OpenAI(base_url='http://localhost:11434/v1', api_key='ollama'). LocalAI exposes exactly the same OpenAI interface — POST /v1/chat/completions, POST /v1/embeddings, POST /v1/images/generations — allowing you to migrate an existing application by only changing the base URL and key. Frameworks such as LangChain, LlamaIndex and Semantic Kernel all have an Ollama or OpenAI connector compatible with both solutions. For RAG, pair either with Weaviate, Qdrant or Chroma: the LLM generates, the vector engine filters relevant contexts, everything stays on your server.
Troubleshooting: Common Errors
OOM error (Out of Memory): the model is too large for available RAM. Check the actual model size with ollama list, reduce quantization (Q4 instead of Q5 or Q8), or choose a smaller model. If the error persists, reduce num_ctx (context window) in the generation options. Model not found: with Ollama, ollama pull can fail if the name is wrong — check the exact listing at ollama.com/library; with LocalAI, verify that the GGUF file is in the /models folder and that its name matches the identifier declared in the configuration. GPU not detected: on a VPS with an NVIDIA GPU, Ollama requires the NVIDIA Container Toolkit runtime installed on the host and the runtime: nvidia parameter in the Compose file; check with ollama ps that the model is actually offloaded to GPU. Timeout on long generations: increase the reverse proxy timeout (Nginx: proxy_read_timeout 300s) and the client-side timeout. For intensive use, prefer streaming ("stream": true) which returns tokens as they are generated and avoids timeouts on long responses.
Ollama or LocalAI: Which Profile Fits Which Tool?
Choose Ollama if you are new to self-hosted LLMs, if you want a developer experience close to a package manager, if your use case is rapid prototyping, an internal chatbot or a code assistant, or if you plan to connect Open WebUI for an immediate user interface. Ollama handles downloads, quantization and CPU/GPU switching on its own — no configuration is required beyond the Compose file. Choose LocalAI if you are migrating an existing application that calls the OpenAI API and want to avoid any code changes on the client side, if you need modalities beyond text (image generation, TTS, STT) in a single container, or if you manage varied model formats (GPTQ, ONNX, exllama2) that the llama.cpp backend alone does not cover. LocalAI demands more initial configuration but offers greater control in production.