Build, host and run AI solutions.

Run open-source LLMs on CPU — no GPU required. OpenAI-compatible API, 200+ models, self-hosted in one command.
LocalAI (MIT, ~47k GitHub stars) is a free, open-source alternative to the OpenAI API. It runs on consumer-grade hardware without a GPU — any VPS with 2 GB RAM can serve Llama 3, Mistral, Phi-3, Qwen and 200+ other models via an OpenAI-compatible REST API. Your existing code that calls `openai.ChatCompletion.create()` works without modification: change the base URL to your VPS and all API calls are routed locally, with zero per-token cost and complete data privacy.
Deployed on a ServOrbit VPS, LocalAI becomes your private LLM backend: an endpoint your team's applications, RAG pipelines and agent frameworks call instead of OpenAI — with no cloud dependency, no usage limits, and no data leaving your infrastructure. LocalAI supports text generation, function calling, image generation (Stable Diffusion), speech-to-text (Whisper), and text embeddings — all from a single container.
curl http://localhost:8080/models/apply -d '{"id":"llama-3.2-3b-instruct:q4_0"}' downloads and activates any supported model.Replace your OpenAI API calls with a LocalAI endpoint on your VPS. Your SaaS, internal tool, or RAG pipeline sends identical JSON with an identical SDK, at zero cost per token. The endpoint is not public — LocalAI listens on the loopback only, because it ships with no authentication whatsoever: code running on the VPS itself calls http://127.0.0.1:8080/v1/chat/completions, and reaching it from another machine takes an SSH tunnel. Ideal for high-volume use cases where per-token API pricing becomes a bottleneck.
For legal, medical or financial applications where data cannot leave your infrastructure, LocalAI runs entirely offline once the model is downloaded. No requests ever reach an external API. The GGUF model files are stored in a Docker volume on your own disk.
Run embedding models (nomic-embed-text, mxbai-embed-large) locally alongside Qdrant on the same VPS to build a complete semantic search or RAG pipeline. No OpenAI embedding API cost, no data leaving your server, no rate limits.
Guide optimized for ServOrbit Cloud VPS.
Order a ServOrbit VPS with at least 2 GB of RAM; the system installed is Ubuntu 24.04. For comfortable multi-user use, or for models from 7 billion parameters upwards, 4 GB is recommended. Core count matters for speed: 4 vCPU noticeably reduce generation latency on models of 3 to 7 billion parameters.
Installing from your client area creates the localai/localai:latest container — an image of about 900 MB — mounts the local-ai volume on /build/models, and publishes the OpenAI-compatible API server on the loopback interface, at 127.0.0.1:8080. No model ships with it: the first request naming a model triggers its download into the persistent volume.
LocalAI opens without any credentials at all: your first move is to go through the model gallery and download one, since the installation provides none. The OpenAI-compatible API answers at the same address and is protected by no key — as long as the service listens only on the loopback interface, only users with SSH access to the VPS can reach it.
A single API call is enough: curl http://127.0.0.1:8080/models/apply -H 'Content-Type: application/json' -d '{"id":"llama-3.2-3b-instruct:q4_0"}'. LocalAI fetches the GGUF file and registers the model. The Q4 quantisation of a 3 billion parameter model fits in 2 GB of RAM; a 7 billion one calls for about 4.
Test the API from the VPS: curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"llama-3.2-3b-instruct:q4_0","messages":[{"role":"user","content":"Hello!"}]}'. The response has exactly the shape of OpenAI's: changing the base URL in your existing code is all it takes to make it work.
From the VPS itself, set base_url='http://127.0.0.1:8080/v1' and api_key='not-needed' in any OpenAI SDK — that is the normal case, since the application consuming LocalAI runs on the same machine. From your own workstation, open an SSH tunnel — ssh -L 8080:127.0.0.1:8080 root@<your-vps-ip> — and point at http://localhost:8080/v1. The port may be reassigned at install time: use the one shown on the app's card in your client area, 8080 being only the catalogue value. In LangChain that gives ChatOpenAI(base_url=..., api_key='x'); in LiteLLM, prefix the model with localai/; in Open WebUI, declare an OpenAI connection with that same base URL. Your existing code needs no other change.
LocalAI is open by default: no key is required. The application knows how to demand one through the LOCALAI_API_KEY environment variable, but the catalogue installation does not set one — the endpoint is therefore unauthenticated, which is precisely why it is never published behind a domain: an inference engine open to the internet is your CPU time and your bill at the disposal of whoever finds it. To expose a public API, publish an authenticating gateway instead — LiteLLM for instance — and keep LocalAI behind it, on the loopback.
Maintaining a project that uses LocalAI? This button lets your readers deploy it on a VPS in one click, without reading Docker documentation.
[](https://servorbit.com/vps-cloud?template=local-ai&utm_source=deploy-badge&utm_medium=referral&utm_campaign=local-ai)<a href="https://servorbit.com/vps-cloud?template=local-ai&utm_source=deploy-badge&utm_medium=referral&utm_campaign=local-ai"><img src="https://servorbit.com/brand/deploy/button.svg" alt="Deploy LocalAI on ServOrbit" height="40"></a>The button points to a VPS order with the template preselected. The image is served from servorbit.com — nothing to host on your side.
Run open-source LLMs locally via a dead-simple API. Pull Llama 3, Mistral, Qwen or DeepSeek in one command — OpenAI-compatible, zero per-token cost.
Artificial IntelligenceSelf-hosted OpenAI-compatible API gateway for 100+ LLMs — route between Ollama, Anthropic, Azure and more from a single endpoint, with per-key rate limits and spend tracking.
Artificial IntelligenceWeb interface to interact with your local or remote LLMs. Your data stays on your infrastructure — no third party involved.
Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.
Message us on WhatsAppopens in a new tab