Deployment guide

How to Host LocalAI on a VPS

Deploy on a VPS Cloud →

Tutorial

How to Host LocalAI on a VPS

Artificial Intelligence8 min read8 steps

LocalAI is an open source, API-compatible replacement for the OpenAI API: chat, embeddings, image and audio generation, all running locally. On your VPS, it becomes a versatile and private AI API that your applications consume without changing their code. This guide covers the full deployment with Docker Compose, loading GGUF models, connecting your applications and troubleshooting the most common errors.

Contents· Why self-host LocalAI on a VPS1/8
  1. 01Why self-host LocalAI on a VPS
  2. 02The concrete benefits of a self-hosted LocalAI
  3. 03Hardware and software requirements
  4. 04Deploy LocalAI with Docker Compose and HTTPS
  5. 05Choosing and loading your models
  6. 06Connecting your applications to LocalAI
  7. 07Troubleshooting: common errors
  8. 08LocalAI vs Ollama vs vLLM: which LLM server to choose?

Why self-host LocalAI on a VPS

Where Ollama focuses on text, LocalAI aims for broad compatibility with the OpenAI API: it exposes /v1/chat/completions, /v1/embeddings, /v1/images/generations and even audio transcription, on a single API. For a team that already has code written against the OpenAI SDK, LocalAI is a near-transparent replacement: you change the base URL and the key, the rest works. On a VPS, you get this versatility without external dependency or per-call cost. This is especially relevant when you need text generation, vectors for RAG and, occasionally, images all at once, without multiplying providers and keys.

The concrete benefits of a self-hosted LocalAI

  • Drop-in replacement for the OpenAI API: no rewriting of the client code.
  • A single API for chat, embeddings, images and audio.
  • Support for multiple backends (llama.cpp, diffusers, whisper) under a unified interface.
  • Open model formats (GGUF) that are downloadable and interchangeable.
  • No usage-based billing: a fixed VPS cost for all types of generation.
  • Data and generations kept entirely on your server.

Hardware and software requirements

Since LocalAI is versatile, its needs depend on the functions enabled. The figures below are measured minimums — a more powerful VPS reduces response times but does not change feasibility.

Text only, CPU, 7B model (e.g. Mistral-7B-GGUF): 8 GB of RAM, 4 vCPU, 30 GB of disk. At this size, a /v1/chat/completions call takes 5 to 15 seconds depending on the quantization (Q4 is the standard trade-off). A 3B model fits in 4 GB of RAM.

Text + images (7B + Stable Diffusion): 16 GB of RAM, GPU recommended, 50 GB of disk. Without a GPU, generating one image takes several minutes on a CPU — prefer a GPU VPS for regular use and the SINGLE_ACTIVE_BACKEND=true option to avoid saturation when mixing chat and images on the same CPU host.

Required software: Docker 24+ and Docker Compose v2 (docker compose without a hyphen), a dedicated subdomain (e.g. ia.your-domain.com), port 443 open, a DNS record pointing to the VPS IP.

Deploy LocalAI with Docker Compose and HTTPS

  1. Prepare the directory structure

    Over SSH: mkdir -p /opt/localai/{models,config} && cd /opt/localai. The models folder will receive GGUF files; config will receive the YAML manifests that describe each model to LocalAI.

  2. Create the docker-compose.yaml file

    Create /opt/localai/docker-compose.yaml with the following content:

    services:
      localai:
        image: localai/localai:latest-cpu
        restart: unless-stopped
        ports:
          - "127.0.0.1:8080:8080"
        volumes:
          - ./models:/build/models
          - ./config:/build/models
        environment:
          - MODELS_PATH=/build/models
          - SINGLE_ACTIVE_BACKEND=true
          - API_KEY=your-secret-key
        healthcheck:
          test: ["CMD", "curl", "-f", "http://localhost:8080/readyz"]
          interval: 30s
          retries: 3

    Replace latest-cpu with latest-gpu-nvidia-cuda-12 if your VPS has an NVIDIA GPU. API_KEY enables token protection on all endpoints — without this variable, the API is accessible without authentication.

  3. Start the container

    Run docker compose up -d, then follow the startup logs: docker compose logs -f localai. Wait for the message LocalAI API is listening on [::]:8080 before continuing.

  4. Download a first model via the gallery

    The LocalAI gallery groups over 200 pre-configured models. To install quantized Mistral-7B-Instruct Q4:

    curl http://127.0.0.1:8080/models/apply \
      -H 'Authorization: Bearer your-secret-key' \
      -d '{"id": "huggingface@thebloke__mistral-7b-instruct-v0.2-gguf/mistral-7b-instruct-v0.2.Q4_K_M.gguf"}'

    Track progress: curl http://127.0.0.1:8080/models/jobs/<job-id>. The GGUF is downloaded into ./models/ and its YAML manifest is created automatically in ./config/.

  5. Add an embeddings model

    For RAG, install all-MiniLM-L6-v2 (84 MB, 256 MB of RAM):

    # config/all-minilm.yaml
    name: all-minilm
    backend: bert-embeddings
    parameters:
      model: all-MiniLM-L6-v2.gguf

    Download the GGUF manually into ./models/ from HuggingFace, then restart: docker compose restart localai. Verify with curl http://127.0.0.1:8080/v1/models.

  6. Enable Whisper for audio transcription

    Add config/whisper.yaml:

    name: whisper-1
    backend: whisper
    parameters:
      model: whisper-base.en

    The whisper-base model (~142 MB) is downloaded automatically on the first call. Test with:

    curl http://127.0.0.1:8080/v1/audio/transcriptions \
      -H 'Authorization: Bearer your-secret-key' \
      -F [email protected] -F model=whisper-1
  7. Configure Nginx reverse proxy and SSL

    Create /etc/nginx/sites-available/localai:

    server {
      listen 443 ssl;
      server_name ia.your-domain.com;
      ssl_certificate /etc/letsencrypt/live/ia.your-domain.com/fullchain.pem;
      ssl_certificate_key /etc/letsencrypt/live/ia.your-domain.com/privkey.pem;
      location / {
        proxy_pass http://127.0.0.1:8080;
        proxy_read_timeout 300s;
        proxy_send_timeout 300s;
      }
    }

    Obtain the certificate with certbot certonly --nginx -d ia.your-domain.com, then nginx -s reload. The high timeout is necessary for long inferences on CPU.

  8. Test the API over HTTPS

    Validate the complete deployment:

    curl https://ia.your-domain.com/v1/chat/completions \
      -H 'Authorization: Bearer your-secret-key' \
      -H 'Content-Type: application/json' \
      -d '{"model": "mistral-7b-instruct-v0.2.Q4_K_M", "messages": [{"role": "user", "content": "Hello"}]}'

    A JSON response with a choices field confirms the deployment is working correctly.

Only enable the backends you need. Loading a chat model, an embedding model and Stable Diffusion simultaneously on a CPU VPS saturates the RAM and collapses performance. Set SINGLE_ACTIVE_BACKEND=true to keep only one backend resident at a time, and reserve image generation for a GPU VPS if it becomes a regular use.

Choosing and loading your models

LocalAI supports several formats, but the GGUF format (produced by llama.cpp) is the recommended standard for CPU: it is quantized (reduces RAM usage), portable and well supported by the gallery.

GGUF vs other formats: GGUF formats are the only ones directly loadable via llama.cpp without prior conversion. Diffusers models (images) and ONNX (lightweight embeddings) have their own backends and are not interchangeable with GGUFs.

Gallery vs manual download: the gallery is the simplest option — it downloads the GGUF and automatically generates the configuration YAML. For a model not listed in the gallery (a fine-tuned model, a custom GGUF), drop the file into ./models/ and create the YAML manually:

name: my-model
backend: llama
parameters:
  model: my-model.Q4_K_M.gguf
context_size: 4096

List loaded models:

curl https://ia.your-domain.com/v1/models \
  -H 'Authorization: Bearer your-secret-key'

Models whose YAML is present in ./config/ appear in the list, even if their GGUF has not yet been downloaded — in that case, the first call triggers the download.

Connecting your applications to LocalAI

Since LocalAI is compatible with the OpenAI API, migrating an existing application comes down to two changes: the base URL and the key.

Python (openai SDK):

import openai

client = openai.OpenAI(
    base_url="https://ia.your-domain.com/v1",
    api_key="your-secret-key"
)

response = client.chat.completions.create(
    model="mistral-7b-instruct-v0.2.Q4_K_M",
    messages=[{"role": "user", "content": "Summarize this text: ..."}]
)
print(response.choices[0].message.content)

Node.js (openai SDK):

import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'https://ia.your-domain.com/v1',
  apiKey: 'your-secret-key'
});

const completion = await client.chat.completions.create({
  model: 'mistral-7b-instruct-v0.2.Q4_K_M',
  messages: [{ role: 'user', content: 'Hello' }]
});
console.log(completion.choices[0].message.content);

Direct curl:

curl https://ia.your-domain.com/v1/embeddings \
  -H 'Authorization: Bearer your-secret-key' \
  -H 'Content-Type: application/json' \
  -d '{"model": "all-minilm", "input": "Text to vectorize"}'

LangChain, LlamaIndex and most RAG frameworks accept an openai_api_base or base_url parameter — the change is identical.

Troubleshooting: common errors

404 model not found or model not loaded: the model name passed in model: does not match any YAML manifest in ./config/. Check with GET /v1/models: the displayed name must be used exactly as shown. If the model does not appear, verify the YAML is in the mounted folder and restart the container.

Very slow inference or killed (OOM): the model exceeds available RAM. A 7B Q4 model requires approximately 6 GB of free RAM at load time. Check with docker stats localai during an inference. Solutions: switch to a lighter quantization (Q3_K_S or Q2_K), reduce the context_size in the YAML, or upgrade the VPS.

SINGLE_ACTIVE_BACKEND error: LocalAI is configured with SINGLE_ACTIVE_BACKEND=true and a request attempts to load a second model while another is active. This is the expected behavior on CPU. Wait for the first inference to finish, or disable the option if your VPS has enough RAM for multiple simultaneous models.

GPU not detected (CUDA device not found): verify you are using the latest-gpu-nvidia-cuda-12 image, that the NVIDIA driver is installed on the host (nvidia-smi), and that the container has GPU access (gpus: all in docker-compose.yaml under the deploy: key). Without all three conditions, LocalAI silently falls back to CPU.

401 Unauthorized: the API_KEY variable is set in docker-compose.yaml but the Authorization: Bearer <key> header is missing or incorrect in the call. All endpoints, including /v1/models, require the header once API_KEY is set.

LocalAI vs Ollama vs vLLM: which LLM server to choose?

Scroll the table

LocalAIOllamavLLM
GPU requiredNo (CPU by default)No (CPU by default)Yes (GPU mandatory)
OpenAI-compatible APIYes (complete)Yes (partial)Yes (complete)
Image generationYes (Stable Diffusion)NoNo
Audio transcriptionYes (Whisper)NoNo
Native embeddingsYesYesYes
Model library200+ GGUF modelsCurated model listHuggingFace models
Throughput (tokens/s, 7B)~10-20 CPU / ~80+ GPU~15-25 CPU / ~80+ GPU~150+ GPU only
RAM (7B Q4 model)~6 GB~6 GB~14 GB (FP16)
Main use caseUnified multi-modal APISimple text inferenceHigh-performance production

A complete and private AI API on a ServOrbit Cloud VPS

With a ServOrbit Cloud VPS and preconfigured Docker, deploy LocalAI as a private alternative to the OpenAI API. Choose a CPU or GPU configuration according to your chat, embeddings or image needs.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab