Why self-host Weaviate on a VPS
Weaviate does more than store vectors. Each object carries its properties (text, numbers, dates, references) and one or more vectors, so a single query returns the relevant documents together with all their metadata. The text2vec-* modules compute embeddings at import time, the generative-* modules pass results to an LLM, and reranker modules reorder the candidates. Everything can be queried over REST, GraphQL or gRPC. Self-hosting on a VPS lets you pick which modules are enabled, plug in your own embedding models (for example through Ollama on the same machine) and keep documents and vectors on a server you administer. For an internal documentation assistant, a catalogue search engine or an agent memory, you control latency, retention, the collection schema and how often backups run. And you scale the VPS resources at the pace of your corpus, rather than at the pace of a usage-based billing grid.
What Weaviate brings in practice
- Built-in vectorization — with a
text2vec-*module, you insert raw text and Weaviate computes the vector itself. - Native hybrid search — a
hybridquery combines the vector score and BM25, weighted by thealphaparameter. - RAG inside the database —
generative-*modules send the retrieved objects to an LLM and return the answer along with the sources. - Multi-tenancy — one isolated shard per customer within the same collection, ideal for a SaaS that separates each account's data.
- Named vectors — several vectors per object (title, content, image), each with its own vectorizer and its own index.
- Index compression — PQ, BQ or SQ quantization to cut the memory used by vectors.
- BSD-3-Clause licence — auditable code you can deploy without depending on a managed service.
Sized requirements for a Weaviate VPS
To test with a few thousand objects, 2 GB of RAM is enough. For real use with vectorization delegated to an API or to Ollama, 4 GB and 2 vCPUs are a comfortable starting point. If you run text2vec-transformers on the same machine, plan for 8 GB or more: CPU inference is slow during bulk imports. On the disk side, count 30 to 60 GB of SSD depending on volume, and keep headroom because Weaviate switches to read-only when disk usage crosses a threshold (DISK_USE_READONLY_PERCENTAGE, 90% by default). Three ports matter: 8080 for REST and GraphQL, 50051 for gRPC (used by the Python client v4 for queries and imports), and 2112 for Prometheus metrics if you enable them. None of them should be exposed raw to the Internet: only an HTTPS reverse proxy, on a subdomain such as vectors.example.com, should face the public. Finally, you need Docker and Docker Compose, already present on VPSs created from the Marketplace.
Which vectorizer to choose
Scroll the table
| Option | Where the model runs | When to choose it |
|---|---|---|
| none + your own vectors | In your application, before import | You already compute embeddings or want full control over the model |
| text2vec-openai | External API, key sent in the X-OpenAI-Api-Key header | Quick start, light VPS, data allowed to leave the server |
| text2vec-ollama | Ollama server on the VPS or the private network | Local embeddings without a dedicated GPU, with a model such as nomic-embed-text |
| text2vec-transformers | Inference container next to Weaviate | Model baked into the image, ideally with a GPU for large imports |
| Several via named vectors | A mix of the options above | Compare two models or vectorize title and content separately |
Deploying Weaviate step by step
Prepare the VPS and DNS
Check Docker with
docker compose version. Create a DNS A record forvectors.example.compointing to the VPS IP, then a/opt/weaviatedirectory that will hold thedocker-compose.yml.Write the weaviate service
Declare the
semitechnologies/weaviate:1.27.2image (also published oncr.weaviate.io). Publish the ports locally only:127.0.0.1:8080:8080and127.0.0.1:50051:50051. Docker bypassesufwrules for published ports, so an8080:8080mapping would remain reachable from the Internet even with the firewall on. Mount a named volume on/var/lib/weaviateand addrestart: unless-stopped.Set the base variables
PERSISTENCE_DATA_PATH=/var/lib/weaviatefor persistence,QUERY_DEFAULTS_LIMIT=25for the default number of results,CLUSTER_HOSTNAME=node1for a stable node name: if it changes between restarts, the node no longer finds its state. AddDEFAULT_VECTORIZER_MODULE=noneor the module of your choice, andENABLE_API_BASED_MODULES=trueto enable the modules that call an API (text2vec-openai,text2vec-ollama,generative-openai…).Enable API key authentication
Set
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED=false, thenAUTHENTICATION_APIKEY_ENABLED=true,AUTHENTICATION_APIKEY_ALLOWED_KEYS=admin-key,read-keyandAUTHENTICATION_APIKEY_USERS=admin,app-reader: keys and users are matched in order. Generate the keys withopenssl rand -hex 32. To restrict the second key to reads, addAUTHORIZATION_ADMINLIST_ENABLED=true,AUTHORIZATION_ADMINLIST_USERS=adminandAUTHORIZATION_ADMINLIST_READONLY_USERS=app-reader.Start and check
Run
docker compose up -d, thencurl -s localhost:8080/v1/.well-known/ready, which answers 200 once the node is ready.curl -s -H "Authorization: Bearer $KEY" localhost:8080/v1/metareturns the version and the list of loaded modules: it is the first check to run when a vectorizer seems to be missing.Put Nginx and TLS in front of port 8080
Configure an Nginx
serverforvectors.example.comwithproxy_pass http://127.0.0.1:8080;, issue the certificate withcertbot --nginx -d vectors.example.comand raiseclient_max_body_sizeif you import batches over REST. TheAuthorizationheader is passed through to Weaviate unchanged.Connect with the Python client v4
Install
pip install -U weaviate-client. On the VPS:client = weaviate.connect_to_local(port=8080, grpc_port=50051, auth_credentials=Auth.api_key(KEY)), withfrom weaviate.classes.init import Auth. From another server, useweaviate.connect_to_custom(...)and specify host, port and TLS for both HTTP and gRPC. Check withclient.is_ready()and finish withclient.close().Create a collection and import
client.collections.create("Article", ...)with its properties and vectorizer, thenarticles = client.collections.get("Article"). Import in batches inside awith articles.batch.dynamic() as batch:block by callingbatch.add_object(properties={...}), and addvector=[...]if the vectorizer isnone. Then checkarticles.batch.failed_objectsto spot rejected objects.Connect the rest of the AI stack
For fully local embeddings and generation, install Ollama on the same VPS and point
text2vec-ollamaandgenerative-ollamaat its API. Open WebUI or your RAG application then query Weaviate as their document memory.
Data modelling: collections, multi-tenancy and named vectors
In Weaviate, data lives in collections (called classes in the older API). A collection defines its typed properties (text, int, number, date, boolean, references to other collections), its vectorizer and its index configuration. Declare properties explicitly instead of letting auto-schema guess them on the first insert: a date field guessed as text can no longer be filtered properly. For a SaaS, enable multi-tenancy at creation time with multi_tenancy_config=Configure.multi_tenancy(enabled=True): each tenant gets its own shard, is queried through collection.with_tenant("customer-a"), and an idle tenant can be deactivated to free memory. Finally, named vectors let you attach several vectors to the same object, for example one for the title and one for the content, each with its own model and index. Be careful: a collection's vectorizer and index type cannot be changed afterwards; a change means a new collection and a re-import.
Querying Weaviate: the useful queries
- Semantic search —
articles.query.near_text(query="contract termination", limit=5)requires a vectorizer; otherwise usenear_vectorwith your own embedding. - Hybrid search —
articles.query.hybrid(query="duplicate invoice", alpha=0.5):alpha=1gives a purely vector search,alpha=0a purely BM25 search. - Keywords only —
articles.query.bm25(query="GDPR")uses the inverted index, useful for references, product codes and proper names. - Filters —
filters=Filter.by_property("language").equal("en"), imported fromweaviate.classes.query, combines with all of the queries above. - Built-in RAG —
articles.generate.near_text(query=..., limit=3, grouped_task="Answer using these excerpts")returns the objects and the answer from the configured LLM. - GraphQL — the
/v1/graphqlendpoint remains available for tools that use it, with thenearText,hybridandbm25operators.
HNSW, flat or dynamic index and memory sizing
By default, each collection uses an HNSW index, fast but held in memory. Weaviate's documentation gives a rule of thumb: plan for about twice the raw size of the vectors. One million 768-dimension vectors in float32 weigh about 3 GB, so around 6 GB of RAM for the index. The flat index keeps nothing in memory and suits small collections, notably the tenants of a SaaS. The dynamic index starts as flat and switches to HNSW above an object threshold; it requires ASYNC_INDEXING=true. To shrink the footprint, enable quantization: PQ (product quantization) compresses heavily after a training phase on your data, BQ (binary quantization) reduces each dimension to one bit and mainly suits high-dimension vectors, SQ (scalar quantization) stores each dimension in one byte. Also set GOMEMLIMIT to about 80% of the RAM allotted to the container so that Go's garbage collector steps in before the kernel kills the process.
Backups and upgrades
Add ENABLE_MODULES=backup-filesystem and BACKUP_FILESYSTEM_PATH=/var/lib/weaviate/backups, then trigger a backup with curl -X POST -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" -d '{"id":"nightly-20261001"}' localhost:8080/v1/backups/filesystem. The status is read from GET /v1/backups/filesystem/nightly-20261001 and restoring goes through POST /v1/backups/filesystem/nightly-20261001/restore. A backup on the same disk does not protect you from losing the VPS: copy the directory elsewhere, or use backup-s3 with BACKUP_S3_BUCKET towards S3-compatible object storage. To upgrade, read the migration notes for the target version, back up, change the image tag, then run docker compose pull && docker compose up -d. Move through minor versions one at a time and never downgrade on the same data without restoring a backup.
Do not leave gRPC port 50051 open to the Internet: applications on the same VPS use it locally, and a remote workstation goes through an ssh -L 50051:localhost:50051 user@vps tunnel. The Marketplace deployment binds this port to 127.0.0.1 and starts with anonymous access open to ease onboarding: enable API key authentication before importing any data. Give your application a read-only key and keep the admin key for imports and backups.
Monitoring Weaviate
Enable PROMETHEUS_MONITORING_ENABLED=true: Weaviate then exposes its metrics on port 2112 (/metrics), to be scraped by Prometheus and displayed in Grafana. Watch process memory, batch import duration and the number of objects per collection first. The /v1/.well-known/live and /v1/.well-known/ready endpoints serve as probes for a Docker healthcheck or a tool such as Uptime Kuma. At system level, docker stats gives the container's live consumption and df -h the volume usage, which deserves close attention since read-only mode is triggered by the disk, not by RAM. Finally, keep an eye on the logs with docker compose logs -f weaviate: module errors and vectorization failures show up there with the name of the object concerned.
Troubleshooting: the most common errors
- 401 response on every request — anonymous access is disabled and the request does not send
Authorization: Bearer <key>, or the key is not listed inAUTHENTICATION_APIKEY_ALLOWED_KEYS. - The Python client v4 refuses to connect — it also checks gRPC: port
50051must be reachable from the client, locally or through an SSH tunnel, andgrpc_portmust match. - Module not found when creating a collection — the module is not loaded: check
ENABLE_MODULES,ENABLE_API_BASED_MODULESand the list returned by/v1/meta. - Container restarting with exit code 137 — the kernel killed it for lack of memory: set
GOMEMLIMIT, enable quantization or add RAM to the VPS. - Imports rejected in read-only mode — the disk crossed
DISK_USE_READONLY_PERCENTAGE: free some space or grow the volume, then set the shards back to writable. - Vectorization errors at import — the provider API key is missing (
X-OpenAI-Api-Keypassed in the clientheaders) or the Ollama server cannot be reached from the Weaviate container.
Weaviate, Qdrant, pgvector or Milvus
Scroll the table
| Criterion | Weaviate | Qdrant | pgvector | Milvus |
|---|---|---|---|---|
| Nature | Vector database in Go | Vector database in Rust | PostgreSQL extension | Distributed vector database |
| Built-in vectorization | Yes, text2vec-* modules | No, embeddings supplied by the client | No, embeddings supplied by the application | Mostly client-side |
| API | REST, GraphQL, gRPC | REST, gRPC | SQL | gRPC, REST, SDKs |
| Hybrid and filters | Native BM25 + vector, filters | Payload filters, sparse vectors | SQL WHERE, PostgreSQL full-text to combine | Scalar filters, BM25 in recent versions |
| Footprint on a VPS | Medium, single container | Low, single container | That of your PostgreSQL | Heavier, etcd and MinIO in standalone mode |
| Compression | PQ, BQ, SQ | Scalar, binary, product | halfvec, bit | Several quantized indexes (IVF_PQ, IVF_SQ8…) |
| Licence | BSD-3-Clause | Apache 2.0 | PostgreSQL Licence | Apache 2.0 |
| Typical use case | All-in-one RAG with built-in vectorization | Lean vector search | Vectors alongside relational data | Very large volumes, cluster deployment |
When Weaviate is the right choice
Weaviate makes sense when you want the database to handle vectorization, hybrid search and generation, rather than assembling them yourself in application code. It fits a documentation assistant, a multilingual catalogue search or a SaaS that isolates each customer in its own tenant. If you already compute your embeddings and want the smallest footprint, Qdrant is leaner. If your data already lives in PostgreSQL and the volume stays modest, pgvector saves you an extra component. Milvus targets very large distributed volumes, at the cost of heavier infrastructure. On a VPS, start with a single node, an HNSW index, an API key and daily backups, then add quantization and multi-tenancy when measurements call for it. To go further, read our guides on Qdrant, Ollama and Open WebUI, which complement Weaviate in a self-hosted RAG stack.
Deploy Weaviate on your VPS
Order a ServOrbit VPS and deploy Weaviate in a few clicks from the Marketplace — Ubuntu 24.04, Docker pre-installed, semitechnologies/weaviate:1.27.2 image ready to run.