Why host LibreChat instead of using ChatGPT Plus
ChatGPT Plus costs $20/month per person, does not allow sharing an API key, sends every message to OpenAI's servers, and locks you into a single interface. Self-hosted LibreChat reverses this: you pay only the tokens you consume from the providers you choose, your conversations stay in your MongoDB, and you can serve ten or a hundred collaborators from a single VPS for under $15/month. API keys are centralized server-side — your users never see them. You choose which models appear in the interface, you can disable certain endpoints for certain roles, and you retain a full log of exchanges for audit or compliance. For organizations concerned about data confidentiality — law firms, HR teams, stealth startups — it is the only solution that provides a real guarantee about where data goes.
ChatGPT Plus vs LibreChat self-hosted
Scroll the table
| Criterion | ChatGPT Plus | LibreChat self-hosted |
|---|---|---|
| Price per user | $20/month | Shared — VPS cost divided by N |
| Available models | GPT-4o, o1 (OpenAI only) | OpenAI, Anthropic, Google, Ollama, Groq… |
| Conversation data | Stored at OpenAI | Stored in your MongoDB |
| API keys | Managed by OpenAI | Yours, never exposed to users |
| Conversation history | At OpenAI, exportable | In your database, freely backupable |
| RAG on your files | Limited (My Files) | Configurable with your own vectorstore |
| Interface customization | None | Complete (presets, visible models, roles) |
| Compliance / audit | Depends on OpenAI terms | Under your full control |
Server requirements
LibreChat relies on several simultaneous services: the main application (Node.js), MongoDB for conversations and users, Meilisearch for history search, and optionally the RAG API for querying documents. For a team of 5 to 20 people using cloud APIs (OpenAI, Anthropic), a 2 vCPU 4 GB RAM VPS is sufficient — around $10 to $15/month at ServOrbit. If you run Ollama on the same server for local models, move to 4 vCPU 8 GB RAM minimum, more if you target 13B-parameter models or larger. Plan 20 GB of disk for MongoDB, Meilisearch indexes, and Docker images. Software-side: Docker Engine 24+ and Docker Compose v2, Git, a subdomain pointing to your VPS (e.g. chat.your-domain.com), and port 443 open in your firewall.
Install LibreChat with Docker Compose
Clone the official repository
SSH into your VPS and clone the repository:
git clone https://github.com/danny-avila/LibreChat && cd LibreChat. The repo already contains a completedocker-compose.ymlthat orchestrates the app, MongoDB, and Meilisearch. Copy the environment example file:cp .env.example .env.Configure the .env file
Open
.envand fill in the essential variables. First generate two random secrets:openssl rand -hex 32twice forJWT_SECRETandCREDS_KEY. Add your API keys:OPENAI_API_KEY=sk-...and/orANTHROPIC_API_KEY=sk-ant-.... To close public registration and only create accounts manually, setALLOW_REGISTRATION=false. To restrict to specific email domains, addALLOWED_DOMAINS=your-company.com. Also setDOMAIN=https://chat.your-domain.com.Launch the stack
Start all services in detached mode:
docker compose up -d. On first launch, Docker downloads the images (~1-2 GB) then starts the containers. Check everything is healthy:docker compose ps— all services should showrunning (healthy). The app listens onlocalhost:3080. Real-time logs are available viadocker compose logs -f app.Create the first admin account
Open
http://localhost:3080in a browser or via an SSH tunnel. Click "Create an account" and sign up with your professional email address. IfALLOW_REGISTRATION=falseis set, temporarily enable registration to create this first account, then lock it again. Once logged in, go to the admin panel to promote your account to admin.Verify LLM connectivity
In the interface, select the
gpt-4omodel (orclaude-3-5-sonnet) and send a test message. If the response arrives correctly, your API key is properly recognized. On a 401 error, check the value in.envthen restart:docker compose restart app.
Configure multiple LLM providers in librechat.yaml
The librechat.yaml file is the heart of multi-provider configuration. It lets you enable or disable endpoints, name the models shown to users, set default parameters (temperature, max tokens), and control access by role. At the project root, copy the template: cp librechat.example.yaml librechat.yaml. Officially supported providers (OpenAI, Anthropic, Google, Azure OpenAI) have a dedicated section under endpoints. For third-party providers or Ollama, use the custom section.
Add OpenAI, Anthropic, and a custom endpoint in librechat.yaml
Enable the OpenAI endpoint
Under the
endpoints.openAIkey, setenabled: trueand list the models to expose:models: { default: [gpt-4o, gpt-4o-mini, o1-mini], fetch: false }. Thefetch: falsefield avoids an API call at startup to fetch the dynamic list — preferable if you want to precisely control what your users see.Enable the Anthropic endpoint
Under
endpoints.anthropic, setenabled: trueand list:models: { default: [claude-3-5-sonnet-20241022, claude-3-haiku-20240307] }. TheANTHROPIC_API_KEYvariable defined in.envis automatically used. You can also settitleConvo: trueso LibreChat automatically generates a conversation title via Claude.Add Google Gemini
Under
endpoints.google, setenabled: trueand fill inGOOGLE_KEYin.envwith your Google AI Studio API key. List the models:gemini-1.5-pro, gemini-1.5-flash. Gemini also supports vision calls (images as input).Declare Groq as a custom endpoint
OpenAI-compatible providers (Groq, Together AI, Perplexity…) are declared under
endpoints.custom. Example for Groq:name: Groq,apiKey: ${GROQ_API_KEY},baseURL: https://api.groq.com/openai/v1,models: { default: [llama-3.1-70b-versatile, mixtral-8x7b-32768] }. AddGROQ_API_KEY=gsk_...to.env, then restart the app.Apply the configuration
After each change to
librechat.yaml, restart only the app service:docker compose restart app. Check in the interface that the new endpoints appear in the model selector. On a YAML parsing error,docker compose logs appwill report it at startup.
Add Ollama for local models (no API cost)
Ollama lets you run open-source models (Llama 3, Mistral, Gemma, Phi-3…) directly on your VPS, without an API key and without data leaving the server. It is the ideal solution for sensitive conversations or to cut costs on low-value requests. On a 4 vCPU 8 GB RAM VPS, 3B to 7B parameter models respond within a few seconds; 13B models are usable but slower. For production use, a VPS with a GPU (or a Hetzner AX41 dedicated server) changes things dramatically.
Deploy Ollama as a Docker service and connect it to LibreChat
Add the Ollama service to docker-compose.yml
Open your
docker-compose.override.yml(create it if it doesn't exist) and add:services: ollama: image: ollama/ollama:latest, volumes: [ollama:/root/.ollama], ports: ["11434:11434"]. Restart the stack:docker compose up -d ollama. Ollama exposes an OpenAI-compatible REST API on port 11434.Download a model
Once the Ollama container is running, download a model:
docker compose exec ollama ollama pull llama3.2:3bfor Llama 3.2 3B (lightweight, ~2 GB), orollama pull mistral:7bfor Mistral 7B (~4 GB). List available models:docker compose exec ollama ollama list.Configure the Ollama endpoint in librechat.yaml
In
librechat.yaml, underendpoints.custom, add:name: Ollama,apiKey: ollama,baseURL: http://ollama:11434/v1(the Docker service name),models: { default: [llama3.2:3b, mistral:7b] },titleConvo: true,titleModel: llama3.2:3b. The service nameollamais resolved automatically by the Docker internal network.Test local inference
In the LibreChat interface, select the "Ollama" endpoint and the
llama3.2:3bmodel. Send a message. The response is generated entirely on your VPS, with no external network call. Note that speed depends on available CPU: on a 4 vCPU VPS, expect 5 to 15 seconds for initial responses.
Authentication and user management
LibreChat offers several levels of authentication, from simple local signup to OAuth2 integration with your identity provider. For a team, the ideal combination is often: closed registration (ALLOW_REGISTRATION=false), accounts created by the admin, and optionally SSO via GitHub, Google, or an in-house OIDC provider (Authentik, Keycloak).
Available authentication options
- Local (email / password): default, works without additional configuration.
- OAuth2 GitHub:
GITHUB_CLIENT_ID+GITHUB_CLIENT_SECRETin.env, then enable underendpoints.social.githubinlibrechat.yaml. - OAuth2 Google:
GOOGLE_CLIENT_ID+GOOGLE_CLIENT_SECRET, same principle. - Generic OIDC: compatible with Authentik, Keycloak, Auth0 — set
OPENID_CLIENT_ID,OPENID_CLIENT_SECRET,OPENID_ISSUER, andOPENID_SCOPEin.env. - Email domain restriction:
ALLOWED_DOMAINS=your-company.comin.env— only addresses from that domain can sign up. - Disable public registration:
ALLOW_REGISTRATION=false— only the admin can create accounts via the administration panel. - Roles: each account is
useroradmin; admins access the user management panel and can modify the endpoints available per user.
RAG with files: query your PDF documents
LibreChat includes an optional RAG (Retrieval-Augmented Generation) service that allows users to upload PDFs, text files, or web pages and query their content in conversations. To enable it, add to .env: RAG_API_URL=http://rag_api:8000. The rag_api service is already present in the official docker-compose.override.yml: launch it with docker compose --profile rag up -d. LibreChat connects to the vectorstore (ChromaDB by default) to index uploaded documents. A 4 GB RAM VPS is sufficient for modest collections (< 500 documents); move to 8 GB for large corpora. Note: the RAG service is CPU-intensive during indexing but lightweight at inference.
Configure nginx as a reverse proxy with HTTPS
Create the nginx vhost
On the server, create
/etc/nginx/sites-available/librechatwith:server { listen 80; server_name chat.your-domain.com; location / { proxy_pass http://localhost:3080; proxy_http_version 1.1; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; proxy_set_header Host $host; proxy_read_timeout 300s; } }. Theproxy_read_timeout 300sis essential to avoid cutting long LLM streaming responses. Enable the site:ln -s /etc/nginx/sites-available/librechat /etc/nginx/sites-enabled/ && nginx -t && systemctl reload nginx.Obtain a Let's Encrypt certificate with Certbot
Install Certbot if needed (
apt install certbot python3-certbot-nginx), then obtain the certificate:certbot --nginx -d chat.your-domain.com. Certbot automatically modifies the vhost to redirect HTTP → HTTPS and sets up automatic renewal. Check renewal:certbot renew --dry-run.Verify WebSocket and streaming
Open
https://chat.your-domain.com, log in, and send a message. Check in DevTools (Network tab) that the connection goes through WebSocket (101 Switching Protocols). If responses appear all at once instead of streaming, it's often an intermediate proxy (Cloudflare) buffering them: enable "Disable buffering" in your Cloudflare configuration or addproxy_buffering offin nginx.Enable automatic updates (optional)
To keep LibreChat up to date without manual intervention, create a weekly cron:
docker compose pull && docker compose up -d --remove-orphans. Note that major updates may modify the MongoDB schema — always read the CHANGELOG before updating in production.
Managing LLM costs with LibreChat
LibreChat does not bill tokens on your behalf — you pay providers directly with your own keys. But self-hosting without monitoring costs can lead to surprises. A few best practices: set token limits per request (maxContextTokens in librechat.yaml) for expensive models like GPT-4o or Claude 3.5 Sonnet; offer gpt-4o-mini or claude-3-haiku as default models for less demanding users; enable Ollama for low-stakes requests (drafts, summaries, rewrites) — actual cost: $0. Monitor your quotas directly in the OpenAI and Anthropic dashboards, and set budget alerts. For a team of 10 people, a mixed budget (GPT-4o for serious cases + Ollama llama3.2 for daily use) typically costs $20 to $80/month in tokens, compared to $200/month for 10 ChatGPT Plus subscriptions.
LibreChat or AnythingLLM: which to choose?
Scroll the table
| Criterion | LibreChat | AnythingLLM |
|---|---|---|
| Main purpose | Multi-provider chat interface | RAG engine over documents |
| Multi-models in one UI | Excellent (OpenAI, Anthropic, Google, local) | Good, but centered on one provider per workspace |
| RAG over documents | Configurable additional module | Core of the product, ready to use |
| Multi-user management | Native accounts and roles | Multi-user mode and partitioned workspaces |
| Deployment architecture | Multi-container stack (Mongo, Meili, RAG) | Single container + vector database |
| Data storage | Self-hosted MongoDB | Volume storage + LanceDB/Qdrant |
| Ideal use case | Team conversational assistant | Queryable knowledge base |
| Resource footprint | Heavier (several services) | Lighter at startup |
Zero-cloud deployment: 100% local
For a fully off-cloud deployment, combine LibreChat with Ollama on the same VPS and disable all cloud endpoints in librechat.yaml. No data will ever leave the server. Add ALLOW_REGISTRATION=false and an in-house OIDC SSO (Authentik) to close the security perimeter. This architecture is particularly suited to law firms, medical teams, or any context subject to GDPR where data cannot transit through American third parties. Total cost is then reduced to the VPS alone — between $10 and $25/month depending on the power chosen at ServOrbit.