Deployment guide

Ollama vs LocalAI: Which Self-Hosted LLM Model Server?

Deploy on a VPS Cloud →

Ollama vs LocalAI: Which Self-Hosted LLM Model Server?

Comparison8 min read6 steps

Running language models on your own server, without sending your prompts to a third party, is the whole point of self-hosted AI. Ollama aims for radical simplicity — one command per model, a clean API, automatic memory management. LocalAI aims for universal compatibility — same endpoints as the OpenAI API, multiple backends, varied formats. Both run on a VPS without a GPU. Here is how to choose, deploy and integrate based on your profile.

Contents· Why Self-Host Your LLM Models on a VPS1/9
  1. 01Why Self-Host Your LLM Models on a VPS
  2. 02The Benefits of a Self-Hosted LLM
  3. 03Prerequisites: Quantized Models and Realistic RAM
  4. 04Deploy Ollama (or LocalAI) on a VPS
  5. 05Managing Models: ollama pull, list and Deletion
  6. 06Performance and Memory: Choosing the Right Quantization
  7. 07Integration: Open WebUI, OpenAI API and SDKs
  8. 08Troubleshooting: Common Errors
  9. 09Ollama or LocalAI: Which Profile Fits Which Tool?

Why Self-Host Your LLM Models on a VPS

Self-hosting a model server means keeping your prompts and sensitive data off commercial APIs, removing per-token billing, and setting the model, its version, and its quantization yourself. On a VPS, you expose a private API to your internal applications (chatbots, RAG, code assistants) without any leak to the outside. Ollama stands out for its radical simplicity: one command to download and run a model, a clean API, automatic memory management. LocalAI positions itself as a 'drop-in' replacement for the OpenAI API: it exposes the same endpoints (chat, embeddings, images, audio) and accepts multiple backends and model formats. The choice comes down to a minimalist experience versus the broadest possible compatibility.

The Benefits of a Self-Hosted LLM

  • Prompts and confidential data that never leave your VPS
  • No per-token billing, predictable cost tied only to the server
  • A private API wired directly into your internal apps
  • Free choice of the model, its size, and its quantization level
  • Compatibility with existing SDKs via an OpenAI-style API
  • Ideal for RAG: pair the LLM with your database and a self-hosted search engine

Prerequisites: Quantized Models and Realistic RAM

Without a GPU, stick to lightweight quantized models. A 3B model in GGUF Q4 runs on a VPS with 4 vCPUs and 8 GB of RAM with acceptable latency for testing; a 7B/8B Q4 model needs 8 to 16 GB of RAM and a good CPU to stay usable. Beyond that, on pure CPU, latency becomes prohibitive: for real-time on large models, a VPS with a GPU is essential. Above all, plan for a generous disk: the weights of several models quickly add up to tens of GB. You will need Docker and Compose, a persistent volume for the models, a subdomain if you expose the API, and a reverse proxy with authentication.

Scroll the table

CriterionOllamaLocalAI
PhilosophySimplicity, one command per modelDrop-in replacement for the OpenAI API
InstallationVery fast, official Docker image readyMore configuration, many backends
OpenAI API compatibilityCompatible `/v1/chat/completions` endpointNative and very complete, all endpoints
Supported modelsOfficial library at ollama.com/libraryGGUF, GPTQ, ONNX, exllama2 and more
GPU supportNVIDIA, AMD, Apple Metal — auto-detectionNVIDIA, AMD, CPU — broad hardware support
Simultaneous multi-modelYes, automatic load/unloadYes, isolated backends per model
Built-in web interfaceNo (Open WebUI separate)No (third-party interface recommended)
Ideal use caseStart fast, prototypeMigrate an OpenAI app to self-hosting

Deploy Ollama (or LocalAI) on a VPS

  1. Provision Storage and the Models Volume

    Create a dedicated volume (/srv/ollama/models) on a sufficiently large disk. Since the weights are large and reusable, they must persist outside the container to avoid re-downloading on every restart.

  2. Launch the Container

    Start ollama/ollama (or localai/localai), mounting the models volume and binding port 11434/8080 locally. On a VPS without a GPU, CPU mode is automatic; with a GPU, enable the appropriate runtime.

  3. Download a Model

    With Ollama, run docker compose exec ollama ollama pull llama3.2:3b. With LocalAI, declare the model in the gallery or drop the GGUF file into the models folder, then verify its loading in the logs.

  4. Test the API

    Make a local call: curl http://127.0.0.1:11434/api/generate for Ollama, or the OpenAI-compatible endpoint POST /v1/chat/completions for LocalAI. Confirm that generation works before any exposure.

  5. Expose Behind an Authenticated Reverse Proxy

    Proxy llm.yourdomain.com to the local port with Caddy or Nginx for TLS, and add an authentication layer (an API key in a header or basic auth). An LLM API open to the Internet is a costly entry point that should never be left uncontrolled.

  6. Connect Your Applications

    Point your existing SDKs to your private base_url. Since LocalAI exposes the OpenAI API, most libraries work by simply changing the URL and the key; on the Ollama side, use its native API or its compatible endpoint.

Managing Models: ollama pull, list and Deletion

With Ollama, all model management goes through the CLI or the REST API. ollama pull llama3.2 downloads the default variant; ollama pull llama3.1:8b-instruct-q4_K_M pulls a precise quantized variant from the official library at ollama.com/library. ollama list shows local models with their disk size and digest. ollama run mistral opens an interactive session in the terminal. ollama rm llama3.2 deletes a model to free up disk space. With LocalAI, the logic differs: models live in the /models folder mounted inside the container. You download a GGUF (llama.cpp format) or GPTQ file and place it in that directory; LocalAI loads it on startup or on demand. The LocalAI gallery offers ready-made configurations for popular models. A practical point: Ollama automatically unloads unused models to free RAM, whereas LocalAI leaves this to your configuration.

Performance and Memory: Choosing the Right Quantization

Quantization reduces the precision of the weights to shrink model size and memory consumption. Q4_K_M and Q5_K_S are the most common formats in practice. Q4_K_M (4-bit, K method, M size) offers an excellent size/quality trade-off for most workloads: a Llama 3.1 8B model in Q4_K_M uses around 4.7 GB of RAM, versus about 16 GB for the non-quantized FP16 version (source: ollama.com/library/llama3.1). Q5_K_S rises to about 5.5 GB but preserves coherence better on long generations. On an 8 GB VPS, target 3B or 7B models in Q4; on 16 GB, a 7B/8B Q5_K_M stays fluid. On pure CPU, expect 3 to 10 tokens/s depending on the model and core count — acceptable for async use, too slow for interactive streaming. With an NVIDIA GPU, Ollama automatically detects CUDA and offloads layers to the GPU; same logic for AMD (ROCm) and Apple Silicon (Metal). LocalAI provides finer control over GPU layer assignment via backend configuration parameters, useful when GPU and CPU must share the workload.

On a VPS without a GPU, the secret to smoothness is quantization: a model in Q4_K_M offers a very favourable size/quality trade-off and fits in RAM where the non-quantized version collapses. Also limit the context window to the strict minimum (for example 4096 tokens): an oversized context multiplies memory consumption and latency with no real benefit for most tasks.

Integration: Open WebUI, OpenAI API and SDKs

Open WebUI is the reference graphical interface for Ollama: it deploys as a single Docker container and connects to Ollama via the OLLAMA_BASE_URL=http://ollama:11434 environment variable. You get a full chat with history, model selection, conversation management — similar to ChatGPT but hosted on your VPS. For programmatic integrations, Ollama exposes /api/generate (simple generation) and /v1/chat/completions (OpenAI-compatible). An existing Python client reconfigures as: client = openai.OpenAI(base_url='http://localhost:11434/v1', api_key='ollama'). LocalAI exposes exactly the same OpenAI interface — POST /v1/chat/completions, POST /v1/embeddings, POST /v1/images/generations — allowing you to migrate an existing application by only changing the base URL and key. Frameworks such as LangChain, LlamaIndex and Semantic Kernel all have an Ollama or OpenAI connector compatible with both solutions. For RAG, pair either with Weaviate, Qdrant or Chroma: the LLM generates, the vector engine filters relevant contexts, everything stays on your server.

Troubleshooting: Common Errors

OOM error (Out of Memory): the model is too large for available RAM. Check the actual model size with ollama list, reduce quantization (Q4 instead of Q5 or Q8), or choose a smaller model. If the error persists, reduce num_ctx (context window) in the generation options. Model not found: with Ollama, ollama pull can fail if the name is wrong — check the exact listing at ollama.com/library; with LocalAI, verify that the GGUF file is in the /models folder and that its name matches the identifier declared in the configuration. GPU not detected: on a VPS with an NVIDIA GPU, Ollama requires the NVIDIA Container Toolkit runtime installed on the host and the runtime: nvidia parameter in the Compose file; check with ollama ps that the model is actually offloaded to GPU. Timeout on long generations: increase the reverse proxy timeout (Nginx: proxy_read_timeout 300s) and the client-side timeout. For intensive use, prefer streaming ("stream": true) which returns tokens as they are generated and avoids timeouts on long responses.

Ollama or LocalAI: Which Profile Fits Which Tool?

Choose Ollama if you are new to self-hosted LLMs, if you want a developer experience close to a package manager, if your use case is rapid prototyping, an internal chatbot or a code assistant, or if you plan to connect Open WebUI for an immediate user interface. Ollama handles downloads, quantization and CPU/GPU switching on its own — no configuration is required beyond the Compose file. Choose LocalAI if you are migrating an existing application that calls the OpenAI API and want to avoid any code changes on the client side, if you need modalities beyond text (image generation, TTS, STT) in a single container, or if you manage varied model formats (GPTQ, ONNX, exllama2) that the llama.cpp backend alone does not cover. LocalAI demands more initial configuration but offers greater control in production.

Your Private LLM Server, Self-Hosted

Deploy Ollama or LocalAI on a ServOrbit Cloud VPS with a Docker template: a private API, reverse proxy, and SSL, to connect your applications without sending your prompts outside.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab