Why self-host LocalAI on a VPS
Where Ollama focuses on text, LocalAI aims for broad compatibility with the OpenAI API: it exposes /v1/chat/completions, /v1/embeddings, /v1/images/generations and even audio transcription, on a single API. For a team that already has code written against the OpenAI SDK, LocalAI is a near-transparent replacement: you change the base URL and the key, the rest works. On a VPS, you get this versatility without external dependency or per-call cost. This is especially relevant when you need text generation, vectors for RAG and, occasionally, images all at once, without multiplying providers and keys.
The concrete benefits of a self-hosted LocalAI
- Drop-in replacement for the OpenAI API: no rewriting of the client code.
- A single API for chat, embeddings, images and audio.
- Support for multiple backends (llama.cpp, diffusers, whisper) under a unified interface.
- Open model formats (GGUF) that are downloadable and interchangeable.
- No usage-based billing: a fixed VPS cost for all types of generation.
- Data and generations kept entirely on your server.
Hardware and software requirements
Since LocalAI is versatile, its needs depend on the functions enabled. The figures below are measured minimums — a more powerful VPS reduces response times but does not change feasibility.
Text only, CPU, 7B model (e.g. Mistral-7B-GGUF): 8 GB of RAM, 4 vCPU, 30 GB of disk. At this size, a /v1/chat/completions call takes 5 to 15 seconds depending on the quantization (Q4 is the standard trade-off). A 3B model fits in 4 GB of RAM.
Text + images (7B + Stable Diffusion): 16 GB of RAM, GPU recommended, 50 GB of disk. Without a GPU, generating one image takes several minutes on a CPU — prefer a GPU VPS for regular use and the SINGLE_ACTIVE_BACKEND=true option to avoid saturation when mixing chat and images on the same CPU host.
Required software: Docker 24+ and Docker Compose v2 (docker compose without a hyphen), a dedicated subdomain (e.g. ia.your-domain.com), port 443 open, a DNS record pointing to the VPS IP.
Deploy LocalAI with Docker Compose and HTTPS
Prepare the directory structure
Over SSH:
mkdir -p /opt/localai/{models,config} && cd /opt/localai. Themodelsfolder will receive GGUF files;configwill receive the YAML manifests that describe each model to LocalAI.Create the docker-compose.yaml file
Create
/opt/localai/docker-compose.yamlwith the following content:services: localai: image: localai/localai:latest-cpu restart: unless-stopped ports: - "127.0.0.1:8080:8080" volumes: - ./models:/build/models - ./config:/build/models environment: - MODELS_PATH=/build/models - SINGLE_ACTIVE_BACKEND=true - API_KEY=your-secret-key healthcheck: test: ["CMD", "curl", "-f", "http://localhost:8080/readyz"] interval: 30s retries: 3Replace
latest-cpuwithlatest-gpu-nvidia-cuda-12if your VPS has an NVIDIA GPU.API_KEYenables token protection on all endpoints — without this variable, the API is accessible without authentication.Start the container
Run
docker compose up -d, then follow the startup logs:docker compose logs -f localai. Wait for the messageLocalAI API is listening on [::]:8080before continuing.Download a first model via the gallery
The LocalAI gallery groups over 200 pre-configured models. To install quantized Mistral-7B-Instruct Q4:
curl http://127.0.0.1:8080/models/apply \ -H 'Authorization: Bearer your-secret-key' \ -d '{"id": "huggingface@thebloke__mistral-7b-instruct-v0.2-gguf/mistral-7b-instruct-v0.2.Q4_K_M.gguf"}'Track progress:
curl http://127.0.0.1:8080/models/jobs/<job-id>. The GGUF is downloaded into./models/and its YAML manifest is created automatically in./config/.Add an embeddings model
For RAG, install
all-MiniLM-L6-v2(84 MB, 256 MB of RAM):# config/all-minilm.yaml name: all-minilm backend: bert-embeddings parameters: model: all-MiniLM-L6-v2.ggufDownload the GGUF manually into
./models/from HuggingFace, then restart:docker compose restart localai. Verify withcurl http://127.0.0.1:8080/v1/models.Enable Whisper for audio transcription
Add
config/whisper.yaml:name: whisper-1 backend: whisper parameters: model: whisper-base.enThe whisper-base model (~142 MB) is downloaded automatically on the first call. Test with:
curl http://127.0.0.1:8080/v1/audio/transcriptions \ -H 'Authorization: Bearer your-secret-key' \ -F [email protected] -F model=whisper-1Configure Nginx reverse proxy and SSL
Create
/etc/nginx/sites-available/localai:server { listen 443 ssl; server_name ia.your-domain.com; ssl_certificate /etc/letsencrypt/live/ia.your-domain.com/fullchain.pem; ssl_certificate_key /etc/letsencrypt/live/ia.your-domain.com/privkey.pem; location / { proxy_pass http://127.0.0.1:8080; proxy_read_timeout 300s; proxy_send_timeout 300s; } }Obtain the certificate with
certbot certonly --nginx -d ia.your-domain.com, thennginx -s reload. The high timeout is necessary for long inferences on CPU.Test the API over HTTPS
Validate the complete deployment:
curl https://ia.your-domain.com/v1/chat/completions \ -H 'Authorization: Bearer your-secret-key' \ -H 'Content-Type: application/json' \ -d '{"model": "mistral-7b-instruct-v0.2.Q4_K_M", "messages": [{"role": "user", "content": "Hello"}]}'A JSON response with a
choicesfield confirms the deployment is working correctly.
Only enable the backends you need. Loading a chat model, an embedding model and Stable Diffusion simultaneously on a CPU VPS saturates the RAM and collapses performance. Set SINGLE_ACTIVE_BACKEND=true to keep only one backend resident at a time, and reserve image generation for a GPU VPS if it becomes a regular use.
Choosing and loading your models
LocalAI supports several formats, but the GGUF format (produced by llama.cpp) is the recommended standard for CPU: it is quantized (reduces RAM usage), portable and well supported by the gallery.
GGUF vs other formats: GGUF formats are the only ones directly loadable via llama.cpp without prior conversion. Diffusers models (images) and ONNX (lightweight embeddings) have their own backends and are not interchangeable with GGUFs.
Gallery vs manual download: the gallery is the simplest option — it downloads the GGUF and automatically generates the configuration YAML. For a model not listed in the gallery (a fine-tuned model, a custom GGUF), drop the file into ./models/ and create the YAML manually:
name: my-model
backend: llama
parameters:
model: my-model.Q4_K_M.gguf
context_size: 4096List loaded models:
curl https://ia.your-domain.com/v1/models \
-H 'Authorization: Bearer your-secret-key'Models whose YAML is present in ./config/ appear in the list, even if their GGUF has not yet been downloaded — in that case, the first call triggers the download.
Connecting your applications to LocalAI
Since LocalAI is compatible with the OpenAI API, migrating an existing application comes down to two changes: the base URL and the key.
Python (openai SDK):
import openai
client = openai.OpenAI(
base_url="https://ia.your-domain.com/v1",
api_key="your-secret-key"
)
response = client.chat.completions.create(
model="mistral-7b-instruct-v0.2.Q4_K_M",
messages=[{"role": "user", "content": "Summarize this text: ..."}]
)
print(response.choices[0].message.content)Node.js (openai SDK):
import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://ia.your-domain.com/v1',
apiKey: 'your-secret-key'
});
const completion = await client.chat.completions.create({
model: 'mistral-7b-instruct-v0.2.Q4_K_M',
messages: [{ role: 'user', content: 'Hello' }]
});
console.log(completion.choices[0].message.content);Direct curl:
curl https://ia.your-domain.com/v1/embeddings \
-H 'Authorization: Bearer your-secret-key' \
-H 'Content-Type: application/json' \
-d '{"model": "all-minilm", "input": "Text to vectorize"}'LangChain, LlamaIndex and most RAG frameworks accept an openai_api_base or base_url parameter — the change is identical.
Troubleshooting: common errors
404 model not found or model not loaded: the model name passed in model: does not match any YAML manifest in ./config/. Check with GET /v1/models: the displayed name must be used exactly as shown. If the model does not appear, verify the YAML is in the mounted folder and restart the container.
Very slow inference or killed (OOM): the model exceeds available RAM. A 7B Q4 model requires approximately 6 GB of free RAM at load time. Check with docker stats localai during an inference. Solutions: switch to a lighter quantization (Q3_K_S or Q2_K), reduce the context_size in the YAML, or upgrade the VPS.
SINGLE_ACTIVE_BACKEND error: LocalAI is configured with SINGLE_ACTIVE_BACKEND=true and a request attempts to load a second model while another is active. This is the expected behavior on CPU. Wait for the first inference to finish, or disable the option if your VPS has enough RAM for multiple simultaneous models.
GPU not detected (CUDA device not found): verify you are using the latest-gpu-nvidia-cuda-12 image, that the NVIDIA driver is installed on the host (nvidia-smi), and that the container has GPU access (gpus: all in docker-compose.yaml under the deploy: key). Without all three conditions, LocalAI silently falls back to CPU.
401 Unauthorized: the API_KEY variable is set in docker-compose.yaml but the Authorization: Bearer <key> header is missing or incorrect in the call. All endpoints, including /v1/models, require the header once API_KEY is set.
LocalAI vs Ollama vs vLLM: which LLM server to choose?
Scroll the table
| LocalAI | Ollama | vLLM | |
|---|---|---|---|
| GPU required | No (CPU by default) | No (CPU by default) | Yes (GPU mandatory) |
| OpenAI-compatible API | Yes (complete) | Yes (partial) | Yes (complete) |
| Image generation | Yes (Stable Diffusion) | No | No |
| Audio transcription | Yes (Whisper) | No | No |
| Native embeddings | Yes | Yes | Yes |
| Model library | 200+ GGUF models | Curated model list | HuggingFace models |
| Throughput (tokens/s, 7B) | ~10-20 CPU / ~80+ GPU | ~15-25 CPU / ~80+ GPU | ~150+ GPU only |
| RAM (7B Q4 model) | ~6 GB | ~6 GB | ~14 GB (FP16) |
| Main use case | Unified multi-modal API | Simple text inference | High-performance production |