Deployment guide

How to Host LibreChat on a VPS

Deploy on a VPS Cloud →

Tutorial

How to Host LibreChat on a VPS

Artificial Intelligence11 min read18 steps

LibreChat reproduces the ChatGPT experience in open source, but goes much further: a unified interface for OpenAI, Anthropic Claude, Google Gemini, local models via Ollama, and many other providers, all behind the same polished UI. Hosted on your own VPS, it gives your team a private, multi-user conversational assistant with history, presets, RAG over your documents, and fine-grained access management — without your conversations transiting through provider servers. No more $20/month per employee for ChatGPT Plus: a single shared API key, all models available, and your data stays in your MongoDB. This guide covers the full installation with Docker Compose, multi-provider configuration in `librechat.yaml`, adding Ollama for models with no API cost, user and authentication management, and setting up the nginx reverse proxy with HTTPS.

Contents· Why host LibreChat instead of using ChatGPT Plus1/15
  1. 01Why host LibreChat instead of using ChatGPT Plus
  2. 02ChatGPT Plus vs LibreChat self-hosted
  3. 03Server requirements
  4. 04Install LibreChat with Docker Compose
  5. 05Configure multiple LLM providers in librechat.yaml
  6. 06Add OpenAI, Anthropic, and a custom endpoint in librechat.yaml
  7. 07Add Ollama for local models (no API cost)
  8. 08Deploy Ollama as a Docker service and connect it to LibreChat
  9. 09Authentication and user management
  10. 10Available authentication options
  11. 11RAG with files: query your PDF documents
  12. 12Configure nginx as a reverse proxy with HTTPS
  13. 13Managing LLM costs with LibreChat
  14. 14LibreChat or AnythingLLM: which to choose?
  15. 15Zero-cloud deployment: 100% local

Why host LibreChat instead of using ChatGPT Plus

ChatGPT Plus costs $20/month per person, does not allow sharing an API key, sends every message to OpenAI's servers, and locks you into a single interface. Self-hosted LibreChat reverses this: you pay only the tokens you consume from the providers you choose, your conversations stay in your MongoDB, and you can serve ten or a hundred collaborators from a single VPS for under $15/month. API keys are centralized server-side — your users never see them. You choose which models appear in the interface, you can disable certain endpoints for certain roles, and you retain a full log of exchanges for audit or compliance. For organizations concerned about data confidentiality — law firms, HR teams, stealth startups — it is the only solution that provides a real guarantee about where data goes.

ChatGPT Plus vs LibreChat self-hosted

Scroll the table

CriterionChatGPT PlusLibreChat self-hosted
Price per user$20/monthShared — VPS cost divided by N
Available modelsGPT-4o, o1 (OpenAI only)OpenAI, Anthropic, Google, Ollama, Groq…
Conversation dataStored at OpenAIStored in your MongoDB
API keysManaged by OpenAIYours, never exposed to users
Conversation historyAt OpenAI, exportableIn your database, freely backupable
RAG on your filesLimited (My Files)Configurable with your own vectorstore
Interface customizationNoneComplete (presets, visible models, roles)
Compliance / auditDepends on OpenAI termsUnder your full control

Server requirements

LibreChat relies on several simultaneous services: the main application (Node.js), MongoDB for conversations and users, Meilisearch for history search, and optionally the RAG API for querying documents. For a team of 5 to 20 people using cloud APIs (OpenAI, Anthropic), a 2 vCPU 4 GB RAM VPS is sufficient — around $10 to $15/month at ServOrbit. If you run Ollama on the same server for local models, move to 4 vCPU 8 GB RAM minimum, more if you target 13B-parameter models or larger. Plan 20 GB of disk for MongoDB, Meilisearch indexes, and Docker images. Software-side: Docker Engine 24+ and Docker Compose v2, Git, a subdomain pointing to your VPS (e.g. chat.your-domain.com), and port 443 open in your firewall.

Install LibreChat with Docker Compose

  1. Clone the official repository

    SSH into your VPS and clone the repository: git clone https://github.com/danny-avila/LibreChat && cd LibreChat. The repo already contains a complete docker-compose.yml that orchestrates the app, MongoDB, and Meilisearch. Copy the environment example file: cp .env.example .env.

  2. Configure the .env file

    Open .env and fill in the essential variables. First generate two random secrets: openssl rand -hex 32 twice for JWT_SECRET and CREDS_KEY. Add your API keys: OPENAI_API_KEY=sk-... and/or ANTHROPIC_API_KEY=sk-ant-.... To close public registration and only create accounts manually, set ALLOW_REGISTRATION=false. To restrict to specific email domains, add ALLOWED_DOMAINS=your-company.com. Also set DOMAIN=https://chat.your-domain.com.

  3. Launch the stack

    Start all services in detached mode: docker compose up -d. On first launch, Docker downloads the images (~1-2 GB) then starts the containers. Check everything is healthy: docker compose ps — all services should show running (healthy). The app listens on localhost:3080. Real-time logs are available via docker compose logs -f app.

  4. Create the first admin account

    Open http://localhost:3080 in a browser or via an SSH tunnel. Click "Create an account" and sign up with your professional email address. If ALLOW_REGISTRATION=false is set, temporarily enable registration to create this first account, then lock it again. Once logged in, go to the admin panel to promote your account to admin.

  5. Verify LLM connectivity

    In the interface, select the gpt-4o model (or claude-3-5-sonnet) and send a test message. If the response arrives correctly, your API key is properly recognized. On a 401 error, check the value in .env then restart: docker compose restart app.

Configure multiple LLM providers in librechat.yaml

The librechat.yaml file is the heart of multi-provider configuration. It lets you enable or disable endpoints, name the models shown to users, set default parameters (temperature, max tokens), and control access by role. At the project root, copy the template: cp librechat.example.yaml librechat.yaml. Officially supported providers (OpenAI, Anthropic, Google, Azure OpenAI) have a dedicated section under endpoints. For third-party providers or Ollama, use the custom section.

Add OpenAI, Anthropic, and a custom endpoint in librechat.yaml

  1. Enable the OpenAI endpoint

    Under the endpoints.openAI key, set enabled: true and list the models to expose: models: { default: [gpt-4o, gpt-4o-mini, o1-mini], fetch: false }. The fetch: false field avoids an API call at startup to fetch the dynamic list — preferable if you want to precisely control what your users see.

  2. Enable the Anthropic endpoint

    Under endpoints.anthropic, set enabled: true and list: models: { default: [claude-3-5-sonnet-20241022, claude-3-haiku-20240307] }. The ANTHROPIC_API_KEY variable defined in .env is automatically used. You can also set titleConvo: true so LibreChat automatically generates a conversation title via Claude.

  3. Add Google Gemini

    Under endpoints.google, set enabled: true and fill in GOOGLE_KEY in .env with your Google AI Studio API key. List the models: gemini-1.5-pro, gemini-1.5-flash. Gemini also supports vision calls (images as input).

  4. Declare Groq as a custom endpoint

    OpenAI-compatible providers (Groq, Together AI, Perplexity…) are declared under endpoints.custom. Example for Groq: name: Groq, apiKey: ${GROQ_API_KEY}, baseURL: https://api.groq.com/openai/v1, models: { default: [llama-3.1-70b-versatile, mixtral-8x7b-32768] }. Add GROQ_API_KEY=gsk_... to .env, then restart the app.

  5. Apply the configuration

    After each change to librechat.yaml, restart only the app service: docker compose restart app. Check in the interface that the new endpoints appear in the model selector. On a YAML parsing error, docker compose logs app will report it at startup.

Add Ollama for local models (no API cost)

Ollama lets you run open-source models (Llama 3, Mistral, Gemma, Phi-3…) directly on your VPS, without an API key and without data leaving the server. It is the ideal solution for sensitive conversations or to cut costs on low-value requests. On a 4 vCPU 8 GB RAM VPS, 3B to 7B parameter models respond within a few seconds; 13B models are usable but slower. For production use, a VPS with a GPU (or a Hetzner AX41 dedicated server) changes things dramatically.

Deploy Ollama as a Docker service and connect it to LibreChat

  1. Add the Ollama service to docker-compose.yml

    Open your docker-compose.override.yml (create it if it doesn't exist) and add: services: ollama: image: ollama/ollama:latest, volumes: [ollama:/root/.ollama], ports: ["11434:11434"]. Restart the stack: docker compose up -d ollama. Ollama exposes an OpenAI-compatible REST API on port 11434.

  2. Download a model

    Once the Ollama container is running, download a model: docker compose exec ollama ollama pull llama3.2:3b for Llama 3.2 3B (lightweight, ~2 GB), or ollama pull mistral:7b for Mistral 7B (~4 GB). List available models: docker compose exec ollama ollama list.

  3. Configure the Ollama endpoint in librechat.yaml

    In librechat.yaml, under endpoints.custom, add: name: Ollama, apiKey: ollama, baseURL: http://ollama:11434/v1 (the Docker service name), models: { default: [llama3.2:3b, mistral:7b] }, titleConvo: true, titleModel: llama3.2:3b. The service name ollama is resolved automatically by the Docker internal network.

  4. Test local inference

    In the LibreChat interface, select the "Ollama" endpoint and the llama3.2:3b model. Send a message. The response is generated entirely on your VPS, with no external network call. Note that speed depends on available CPU: on a 4 vCPU VPS, expect 5 to 15 seconds for initial responses.

Authentication and user management

LibreChat offers several levels of authentication, from simple local signup to OAuth2 integration with your identity provider. For a team, the ideal combination is often: closed registration (ALLOW_REGISTRATION=false), accounts created by the admin, and optionally SSO via GitHub, Google, or an in-house OIDC provider (Authentik, Keycloak).

Available authentication options

  • Local (email / password): default, works without additional configuration.
  • OAuth2 GitHub: GITHUB_CLIENT_ID + GITHUB_CLIENT_SECRET in .env, then enable under endpoints.social.github in librechat.yaml.
  • OAuth2 Google: GOOGLE_CLIENT_ID + GOOGLE_CLIENT_SECRET, same principle.
  • Generic OIDC: compatible with Authentik, Keycloak, Auth0 — set OPENID_CLIENT_ID, OPENID_CLIENT_SECRET, OPENID_ISSUER, and OPENID_SCOPE in .env.
  • Email domain restriction: ALLOWED_DOMAINS=your-company.com in .env — only addresses from that domain can sign up.
  • Disable public registration: ALLOW_REGISTRATION=false — only the admin can create accounts via the administration panel.
  • Roles: each account is user or admin; admins access the user management panel and can modify the endpoints available per user.

RAG with files: query your PDF documents

LibreChat includes an optional RAG (Retrieval-Augmented Generation) service that allows users to upload PDFs, text files, or web pages and query their content in conversations. To enable it, add to .env: RAG_API_URL=http://rag_api:8000. The rag_api service is already present in the official docker-compose.override.yml: launch it with docker compose --profile rag up -d. LibreChat connects to the vectorstore (ChromaDB by default) to index uploaded documents. A 4 GB RAM VPS is sufficient for modest collections (< 500 documents); move to 8 GB for large corpora. Note: the RAG service is CPU-intensive during indexing but lightweight at inference.

Configure nginx as a reverse proxy with HTTPS

  1. Create the nginx vhost

    On the server, create /etc/nginx/sites-available/librechat with: server { listen 80; server_name chat.your-domain.com; location / { proxy_pass http://localhost:3080; proxy_http_version 1.1; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; proxy_set_header Host $host; proxy_read_timeout 300s; } }. The proxy_read_timeout 300s is essential to avoid cutting long LLM streaming responses. Enable the site: ln -s /etc/nginx/sites-available/librechat /etc/nginx/sites-enabled/ && nginx -t && systemctl reload nginx.

  2. Obtain a Let's Encrypt certificate with Certbot

    Install Certbot if needed (apt install certbot python3-certbot-nginx), then obtain the certificate: certbot --nginx -d chat.your-domain.com. Certbot automatically modifies the vhost to redirect HTTP → HTTPS and sets up automatic renewal. Check renewal: certbot renew --dry-run.

  3. Verify WebSocket and streaming

    Open https://chat.your-domain.com, log in, and send a message. Check in DevTools (Network tab) that the connection goes through WebSocket (101 Switching Protocols). If responses appear all at once instead of streaming, it's often an intermediate proxy (Cloudflare) buffering them: enable "Disable buffering" in your Cloudflare configuration or add proxy_buffering off in nginx.

  4. Enable automatic updates (optional)

    To keep LibreChat up to date without manual intervention, create a weekly cron: docker compose pull && docker compose up -d --remove-orphans. Note that major updates may modify the MongoDB schema — always read the CHANGELOG before updating in production.

Managing LLM costs with LibreChat

LibreChat does not bill tokens on your behalf — you pay providers directly with your own keys. But self-hosting without monitoring costs can lead to surprises. A few best practices: set token limits per request (maxContextTokens in librechat.yaml) for expensive models like GPT-4o or Claude 3.5 Sonnet; offer gpt-4o-mini or claude-3-haiku as default models for less demanding users; enable Ollama for low-stakes requests (drafts, summaries, rewrites) — actual cost: $0. Monitor your quotas directly in the OpenAI and Anthropic dashboards, and set budget alerts. For a team of 10 people, a mixed budget (GPT-4o for serious cases + Ollama llama3.2 for daily use) typically costs $20 to $80/month in tokens, compared to $200/month for 10 ChatGPT Plus subscriptions.

LibreChat or AnythingLLM: which to choose?

Scroll the table

CriterionLibreChatAnythingLLM
Main purposeMulti-provider chat interfaceRAG engine over documents
Multi-models in one UIExcellent (OpenAI, Anthropic, Google, local)Good, but centered on one provider per workspace
RAG over documentsConfigurable additional moduleCore of the product, ready to use
Multi-user managementNative accounts and rolesMulti-user mode and partitioned workspaces
Deployment architectureMulti-container stack (Mongo, Meili, RAG)Single container + vector database
Data storageSelf-hosted MongoDBVolume storage + LanceDB/Qdrant
Ideal use caseTeam conversational assistantQueryable knowledge base
Resource footprintHeavier (several services)Lighter at startup

Zero-cloud deployment: 100% local

For a fully off-cloud deployment, combine LibreChat with Ollama on the same VPS and disable all cloud endpoints in librechat.yaml. No data will ever leave the server. Add ALLOW_REGISTRATION=false and an in-house OIDC SSO (Authentik) to close the security perimeter. This architecture is particularly suited to law firms, medical teams, or any context subject to GDPR where data cannot transit through American third parties. Total cost is then reduced to the VPS alone — between $10 and $25/month depending on the power chosen at ServOrbit.

Install LibreChat on a ServOrbit VPS with one click

LibreChat is available in the ServOrbit Marketplace. Order a Cloud VPS and the multi-model interface installs automatically, MongoDB included.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab