Deployment guide

Hosting Weaviate on a VPS

Deploy on a VPS Cloud →

Tutorial

Hosting Weaviate on a VPS

Databases12 min read9 steps

Weaviate is an open-source vector database written in Go, built for semantic search and RAG. What sets it apart: modules that vectorize your objects at import time and can generate an answer from the results, with no external pipeline. This guide deploys Weaviate on a VPS with Docker Compose, secures the API with keys, explains how to choose a vectorizer, how to size memory for the HNSW index, how to back up, and how to fix the most common errors. A comparison with Qdrant, pgvector and Milvus helps you decide.

Contents· Why self-host Weaviate on a VPS1/14
  1. 01Why self-host Weaviate on a VPS
  2. 02What Weaviate brings in practice
  3. 03Sized requirements for a Weaviate VPS
  4. 04Which vectorizer to choose
  5. 05Deploying Weaviate step by step
  6. 06Data modelling: collections, multi-tenancy and named vectors
  7. 07Querying Weaviate: the useful queries
  8. 08HNSW, flat or dynamic index and memory sizing
  9. 09Backups and upgrades
  10. 10Monitoring Weaviate
  11. 11Troubleshooting: the most common errors
  12. 12Weaviate, Qdrant, pgvector or Milvus
  13. 13When Weaviate is the right choice
  14. 14Deploy Weaviate on your VPS

Why self-host Weaviate on a VPS

Weaviate does more than store vectors. Each object carries its properties (text, numbers, dates, references) and one or more vectors, so a single query returns the relevant documents together with all their metadata. The text2vec-* modules compute embeddings at import time, the generative-* modules pass results to an LLM, and reranker modules reorder the candidates. Everything can be queried over REST, GraphQL or gRPC. Self-hosting on a VPS lets you pick which modules are enabled, plug in your own embedding models (for example through Ollama on the same machine) and keep documents and vectors on a server you administer. For an internal documentation assistant, a catalogue search engine or an agent memory, you control latency, retention, the collection schema and how often backups run. And you scale the VPS resources at the pace of your corpus, rather than at the pace of a usage-based billing grid.

What Weaviate brings in practice

  • Built-in vectorization — with a text2vec-* module, you insert raw text and Weaviate computes the vector itself.
  • Native hybrid search — a hybrid query combines the vector score and BM25, weighted by the alpha parameter.
  • RAG inside the database — generative-* modules send the retrieved objects to an LLM and return the answer along with the sources.
  • Multi-tenancy — one isolated shard per customer within the same collection, ideal for a SaaS that separates each account's data.
  • Named vectors — several vectors per object (title, content, image), each with its own vectorizer and its own index.
  • Index compression — PQ, BQ or SQ quantization to cut the memory used by vectors.
  • BSD-3-Clause licence — auditable code you can deploy without depending on a managed service.

Sized requirements for a Weaviate VPS

To test with a few thousand objects, 2 GB of RAM is enough. For real use with vectorization delegated to an API or to Ollama, 4 GB and 2 vCPUs are a comfortable starting point. If you run text2vec-transformers on the same machine, plan for 8 GB or more: CPU inference is slow during bulk imports. On the disk side, count 30 to 60 GB of SSD depending on volume, and keep headroom because Weaviate switches to read-only when disk usage crosses a threshold (DISK_USE_READONLY_PERCENTAGE, 90% by default). Three ports matter: 8080 for REST and GraphQL, 50051 for gRPC (used by the Python client v4 for queries and imports), and 2112 for Prometheus metrics if you enable them. None of them should be exposed raw to the Internet: only an HTTPS reverse proxy, on a subdomain such as vectors.example.com, should face the public. Finally, you need Docker and Docker Compose, already present on VPSs created from the Marketplace.

Which vectorizer to choose

Scroll the table

OptionWhere the model runsWhen to choose it
none + your own vectorsIn your application, before importYou already compute embeddings or want full control over the model
text2vec-openaiExternal API, key sent in the X-OpenAI-Api-Key headerQuick start, light VPS, data allowed to leave the server
text2vec-ollamaOllama server on the VPS or the private networkLocal embeddings without a dedicated GPU, with a model such as nomic-embed-text
text2vec-transformersInference container next to WeaviateModel baked into the image, ideally with a GPU for large imports
Several via named vectorsA mix of the options aboveCompare two models or vectorize title and content separately

Deploying Weaviate step by step

  1. Prepare the VPS and DNS

    Check Docker with docker compose version. Create a DNS A record for vectors.example.com pointing to the VPS IP, then a /opt/weaviate directory that will hold the docker-compose.yml.

  2. Write the weaviate service

    Declare the semitechnologies/weaviate:1.27.2 image (also published on cr.weaviate.io). Publish the ports locally only: 127.0.0.1:8080:8080 and 127.0.0.1:50051:50051. Docker bypasses ufw rules for published ports, so an 8080:8080 mapping would remain reachable from the Internet even with the firewall on. Mount a named volume on /var/lib/weaviate and add restart: unless-stopped.

  3. Set the base variables

    PERSISTENCE_DATA_PATH=/var/lib/weaviate for persistence, QUERY_DEFAULTS_LIMIT=25 for the default number of results, CLUSTER_HOSTNAME=node1 for a stable node name: if it changes between restarts, the node no longer finds its state. Add DEFAULT_VECTORIZER_MODULE=none or the module of your choice, and ENABLE_API_BASED_MODULES=true to enable the modules that call an API (text2vec-openai, text2vec-ollama, generative-openai…).

  4. Enable API key authentication

    Set AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED=false, then AUTHENTICATION_APIKEY_ENABLED=true, AUTHENTICATION_APIKEY_ALLOWED_KEYS=admin-key,read-key and AUTHENTICATION_APIKEY_USERS=admin,app-reader: keys and users are matched in order. Generate the keys with openssl rand -hex 32. To restrict the second key to reads, add AUTHORIZATION_ADMINLIST_ENABLED=true, AUTHORIZATION_ADMINLIST_USERS=admin and AUTHORIZATION_ADMINLIST_READONLY_USERS=app-reader.

  5. Start and check

    Run docker compose up -d, then curl -s localhost:8080/v1/.well-known/ready, which answers 200 once the node is ready. curl -s -H "Authorization: Bearer $KEY" localhost:8080/v1/meta returns the version and the list of loaded modules: it is the first check to run when a vectorizer seems to be missing.

  6. Put Nginx and TLS in front of port 8080

    Configure an Nginx server for vectors.example.com with proxy_pass http://127.0.0.1:8080;, issue the certificate with certbot --nginx -d vectors.example.com and raise client_max_body_size if you import batches over REST. The Authorization header is passed through to Weaviate unchanged.

  7. Connect with the Python client v4

    Install pip install -U weaviate-client. On the VPS: client = weaviate.connect_to_local(port=8080, grpc_port=50051, auth_credentials=Auth.api_key(KEY)), with from weaviate.classes.init import Auth. From another server, use weaviate.connect_to_custom(...) and specify host, port and TLS for both HTTP and gRPC. Check with client.is_ready() and finish with client.close().

  8. Create a collection and import

    client.collections.create("Article", ...) with its properties and vectorizer, then articles = client.collections.get("Article"). Import in batches inside a with articles.batch.dynamic() as batch: block by calling batch.add_object(properties={...}), and add vector=[...] if the vectorizer is none. Then check articles.batch.failed_objects to spot rejected objects.

  9. Connect the rest of the AI stack

    For fully local embeddings and generation, install Ollama on the same VPS and point text2vec-ollama and generative-ollama at its API. Open WebUI or your RAG application then query Weaviate as their document memory.

Data modelling: collections, multi-tenancy and named vectors

In Weaviate, data lives in collections (called classes in the older API). A collection defines its typed properties (text, int, number, date, boolean, references to other collections), its vectorizer and its index configuration. Declare properties explicitly instead of letting auto-schema guess them on the first insert: a date field guessed as text can no longer be filtered properly. For a SaaS, enable multi-tenancy at creation time with multi_tenancy_config=Configure.multi_tenancy(enabled=True): each tenant gets its own shard, is queried through collection.with_tenant("customer-a"), and an idle tenant can be deactivated to free memory. Finally, named vectors let you attach several vectors to the same object, for example one for the title and one for the content, each with its own model and index. Be careful: a collection's vectorizer and index type cannot be changed afterwards; a change means a new collection and a re-import.

Querying Weaviate: the useful queries

  • Semantic search — articles.query.near_text(query="contract termination", limit=5) requires a vectorizer; otherwise use near_vector with your own embedding.
  • Hybrid search — articles.query.hybrid(query="duplicate invoice", alpha=0.5): alpha=1 gives a purely vector search, alpha=0 a purely BM25 search.
  • Keywords only — articles.query.bm25(query="GDPR") uses the inverted index, useful for references, product codes and proper names.
  • Filters — filters=Filter.by_property("language").equal("en"), imported from weaviate.classes.query, combines with all of the queries above.
  • Built-in RAG — articles.generate.near_text(query=..., limit=3, grouped_task="Answer using these excerpts") returns the objects and the answer from the configured LLM.
  • GraphQL — the /v1/graphql endpoint remains available for tools that use it, with the nearText, hybrid and bm25 operators.

HNSW, flat or dynamic index and memory sizing

By default, each collection uses an HNSW index, fast but held in memory. Weaviate's documentation gives a rule of thumb: plan for about twice the raw size of the vectors. One million 768-dimension vectors in float32 weigh about 3 GB, so around 6 GB of RAM for the index. The flat index keeps nothing in memory and suits small collections, notably the tenants of a SaaS. The dynamic index starts as flat and switches to HNSW above an object threshold; it requires ASYNC_INDEXING=true. To shrink the footprint, enable quantization: PQ (product quantization) compresses heavily after a training phase on your data, BQ (binary quantization) reduces each dimension to one bit and mainly suits high-dimension vectors, SQ (scalar quantization) stores each dimension in one byte. Also set GOMEMLIMIT to about 80% of the RAM allotted to the container so that Go's garbage collector steps in before the kernel kills the process.

Backups and upgrades

Add ENABLE_MODULES=backup-filesystem and BACKUP_FILESYSTEM_PATH=/var/lib/weaviate/backups, then trigger a backup with curl -X POST -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" -d '{"id":"nightly-20261001"}' localhost:8080/v1/backups/filesystem. The status is read from GET /v1/backups/filesystem/nightly-20261001 and restoring goes through POST /v1/backups/filesystem/nightly-20261001/restore. A backup on the same disk does not protect you from losing the VPS: copy the directory elsewhere, or use backup-s3 with BACKUP_S3_BUCKET towards S3-compatible object storage. To upgrade, read the migration notes for the target version, back up, change the image tag, then run docker compose pull && docker compose up -d. Move through minor versions one at a time and never downgrade on the same data without restoring a backup.

Do not leave gRPC port 50051 open to the Internet: applications on the same VPS use it locally, and a remote workstation goes through an ssh -L 50051:localhost:50051 user@vps tunnel. The Marketplace deployment binds this port to 127.0.0.1 and starts with anonymous access open to ease onboarding: enable API key authentication before importing any data. Give your application a read-only key and keep the admin key for imports and backups.

Monitoring Weaviate

Enable PROMETHEUS_MONITORING_ENABLED=true: Weaviate then exposes its metrics on port 2112 (/metrics), to be scraped by Prometheus and displayed in Grafana. Watch process memory, batch import duration and the number of objects per collection first. The /v1/.well-known/live and /v1/.well-known/ready endpoints serve as probes for a Docker healthcheck or a tool such as Uptime Kuma. At system level, docker stats gives the container's live consumption and df -h the volume usage, which deserves close attention since read-only mode is triggered by the disk, not by RAM. Finally, keep an eye on the logs with docker compose logs -f weaviate: module errors and vectorization failures show up there with the name of the object concerned.

Troubleshooting: the most common errors

  • 401 response on every request — anonymous access is disabled and the request does not send Authorization: Bearer <key>, or the key is not listed in AUTHENTICATION_APIKEY_ALLOWED_KEYS.
  • The Python client v4 refuses to connect — it also checks gRPC: port 50051 must be reachable from the client, locally or through an SSH tunnel, and grpc_port must match.
  • Module not found when creating a collection — the module is not loaded: check ENABLE_MODULES, ENABLE_API_BASED_MODULES and the list returned by /v1/meta.
  • Container restarting with exit code 137 — the kernel killed it for lack of memory: set GOMEMLIMIT, enable quantization or add RAM to the VPS.
  • Imports rejected in read-only mode — the disk crossed DISK_USE_READONLY_PERCENTAGE: free some space or grow the volume, then set the shards back to writable.
  • Vectorization errors at import — the provider API key is missing (X-OpenAI-Api-Key passed in the client headers) or the Ollama server cannot be reached from the Weaviate container.

Weaviate, Qdrant, pgvector or Milvus

Scroll the table

CriterionWeaviateQdrantpgvectorMilvus
NatureVector database in GoVector database in RustPostgreSQL extensionDistributed vector database
Built-in vectorizationYes, text2vec-* modulesNo, embeddings supplied by the clientNo, embeddings supplied by the applicationMostly client-side
APIREST, GraphQL, gRPCREST, gRPCSQLgRPC, REST, SDKs
Hybrid and filtersNative BM25 + vector, filtersPayload filters, sparse vectorsSQL WHERE, PostgreSQL full-text to combineScalar filters, BM25 in recent versions
Footprint on a VPSMedium, single containerLow, single containerThat of your PostgreSQLHeavier, etcd and MinIO in standalone mode
CompressionPQ, BQ, SQScalar, binary, producthalfvec, bitSeveral quantized indexes (IVF_PQ, IVF_SQ8…)
LicenceBSD-3-ClauseApache 2.0PostgreSQL LicenceApache 2.0
Typical use caseAll-in-one RAG with built-in vectorizationLean vector searchVectors alongside relational dataVery large volumes, cluster deployment

When Weaviate is the right choice

Weaviate makes sense when you want the database to handle vectorization, hybrid search and generation, rather than assembling them yourself in application code. It fits a documentation assistant, a multilingual catalogue search or a SaaS that isolates each customer in its own tenant. If you already compute your embeddings and want the smallest footprint, Qdrant is leaner. If your data already lives in PostgreSQL and the volume stays modest, pgvector saves you an extra component. Milvus targets very large distributed volumes, at the cost of heavier infrastructure. On a VPS, start with a single node, an HNSW index, an API key and daily backups, then add quantization and multi-tenancy when measurements call for it. To go further, read our guides on Qdrant, Ollama and Open WebUI, which complement Weaviate in a self-hosted RAG stack.

Deploy Weaviate on your VPS

Order a ServOrbit VPS and deploy Weaviate in a few clicks from the Marketplace — Ubuntu 24.04, Docker pre-installed, semitechnologies/weaviate:1.27.2 image ready to run.

Deploy your Weaviate RAG platform

The ServOrbit Cloud VPS gives you the resources and the Docker environment to host Weaviate and its modules, and to build a complete, secure and sovereign semantic search pipeline.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab