Why self-host Qdrant on a VPS
Qdrant stores and queries embedding vectors to find the most semantically similar items: it is the search building block at the heart of RAG chatbots, recommendation engines and natural-language search. Written in Rust, it is fast and resource-efficient, making it ideal to host on a VPS rather than paying for a managed vector service billed per stored vector.
Self-hosting is essential for AI applications handling sensitive data: your embeddings, often derived from confidential internal documents, never leave your infrastructure. You also control latency, which is decisive when every user request triggers a vector search, and you avoid the rate limits of third-party APIs during load spikes.
Prerequisites and recommended VPS resources
- RAM: minimum 1 GB for testing, 2 GB for a few hundred thousand vectors, 4–8 GB beyond a million high-dimensional vectors (the HNSW index lives in memory).
- Storage: 20–50 GB of NVMe SSD to persist collections and snapshots.
- CPU: 1–2 vCPUs are sufficient for moderate traffic; add cores if you index continuously.
- Docker and Docker Compose installed on the VPS.
- Ports 6333 (REST) and 6334 (gRPC) accessible from your application; 80/443 for the public reverse proxy.
- A DNS A record (e.g.
vectors.mydomain.com) if you expose the API over HTTPS.
Install Qdrant with Docker Compose
Create the directory and compose.yml
On the VPS, create a dedicated folder:
mkdir -p /opt/qdrant && cd /opt/qdrant. Create the followingcompose.yml:services: qdrant: image: qdrant/qdrant:latest restart: unless-stopped ports: - "127.0.0.1:6333:6333" - "127.0.0.1:6334:6334" volumes: - ./qdrant_storage:/qdrant/storage environment: - QDRANT__SERVICE__API_KEY=${QDRANT_API_KEY}Binding ports to
127.0.0.1prevents any direct exposure on the Internet; only the reverse proxy can reach them.Generate the API key and start the container
Create the
.envfile with your key:echo "QDRANT_API_KEY=$(openssl rand -hex 32)" > .envStart Qdrant:
docker compose up -dCheck instance health
Wait a few seconds then verify the health endpoint:
curl -s http://localhost:6333/healthz # Expected: {"title":"qdrant - vector search engine","version":"..."}Create your first collection
Create a test collection with 384-dimensional vectors (typical format for
all-MiniLM-L6-v2):curl -s -X PUT http://localhost:6333/collections/test \ -H 'Content-Type: application/json' \ -H "api-key: $(grep QDRANT_API_KEY .env | cut -d= -f2)" \ -d '{"vectors":{"size":384,"distance":"Cosine"}}'
Securing access: API key and TLS
Qdrant enables no authentication by default: an instance exposed without an API key would allow anyone to read, modify or delete all your collections. The QDRANT__SERVICE__API_KEY environment variable is the first line of defence; the TLS reverse proxy is the second.
For Caddy, the minimal configuration is:
vectors.mydomain.com {
reverse_proxy localhost:6333
}Caddy automatically obtains and renews the Let's Encrypt certificate. For nginx, add inside your server block:
location / {
proxy_pass http://127.0.0.1:6333;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}Enable TLS with Certbot: certbot --nginx -d vectors.mydomain.com.
Test authentication and the reverse proxy
Verify that unauthenticated access is rejected
curl -s https://vectors.mydomain.com/collections # Expected: {"status":{"error":"Unauthorized",...}}Authenticated call with the API key
curl -s https://vectors.mydomain.com/collections \ -H "api-key: YOUR_KEY" # Expected: {"result":{"collections":[]},"status":"ok",...}Verify the TLS certificate
curl -sv https://vectors.mydomain.com/healthz 2>&1 | grep 'SSL connection\|issuer' # Expected: lines confirming Let's Encrypt and TLS 1.3
Create a collection and insert vectors with the Python v1 client
Install the official client
pip install qdrant-client>=1.7Connect with the API key
from qdrant_client import QdrantClient client = QdrantClient( url="https://vectors.mydomain.com", api_key="YOUR_KEY", )Create a collection with VectorParams
from qdrant_client.models import Distance, VectorParams client.recreate_collection( collection_name="documents", vectors_config=VectorParams(size=384, distance=Distance.COSINE), )sizemust match the embedding dimension of your model.Distance.COSINEsuits most sentence-transformer models.Insert points with payload and query
from qdrant_client.models import PointStruct # Insert vectors client.upsert( collection_name="documents", points=[ PointStruct( id=1, vector=[0.1, 0.2, ...], # vector of size 384 payload={"title": "Qdrant Guide", "lang": "en"}, ), ], ) # Search for the 5 nearest neighbours results = client.search( collection_name="documents", query_vector=[0.1, 0.21, ...], limit=5, with_payload=True, ) for r in results: print(r.score, r.payload)
Optimise performance: HNSW tuning
The HNSW (Hierarchical Navigable Small World) algorithm is the core of fast search in Qdrant. Three parameters govern the speed/accuracy/memory trade-off.
Key HNSW parameters
Scroll the table
| Parameter | Default | Recommended range | Impact |
|---|---|---|---|
| m | 16 | 16–64 | Connections per node — higher = better recall, more RAM |
| ef_construct | 100 | 100–400 | Index build quality — higher = better index, slower build |
| ef | 128 (at query) | 64–512 | Search quality — higher = better recall, slower query |
Configure HNSW at collection creation
Pass HNSW parameters directly in recreate_collection:
from qdrant_client.models import HnswConfigDiff
client.recreate_collection(
collection_name="documents",
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
hnsw_config=HnswConfigDiff(m=32, ef_construct=200),
)For balanced search on a corpus of 500,000 vectors, m=32 and ef_construct=200 provide good recall (>0.97) without overloading RAM. Increase ef at query time for cases where precision matters more than latency.
Reduce memory footprint: scalar and binary quantization
- Scalar quantization (uint8): compresses each float32 coordinate (4 bytes) to uint8 (1 byte), roughly a 4× RAM reduction. Slight precision loss offset by oversampling. Enable with
ScalarQuantizationConfig(type=ScalarType.INT8, quantile=0.99, always_ram=True). - Binary quantization: up to 32× reduction by binarising each dimension. Ideal for models trained for binary quantization (e.g. Cohere
embed-english-v3.0). Faster but lower recall without oversampling. on_diskparameter: store the original vectors on disk (on_disk=True) and keep only the quantised index in RAM. Essential on memory-constrained VPS.- Oversampling: with
oversampling=2.0at query time, Qdrant retrieves 2× more candidates from the quantised index then re-ranks them with the original vectors — recall close to full-precision at moderate cost.
Back up and restore with Qdrant snapshots
Create a snapshot of a collection
curl -s -X POST https://vectors.mydomain.com/collections/documents/snapshots \ -H "api-key: YOUR_KEY" # Returns: {"result":{"name":"documents-...","creation_time":"..."}}List available snapshots
curl -s https://vectors.mydomain.com/collections/documents/snapshots \ -H "api-key: YOUR_KEY"Download the snapshot
curl -O https://vectors.mydomain.com/collections/documents/snapshots/SNAPSHOT_NAME \ -H "api-key: YOUR_KEY"Then copy the
.snapshotfile to object storage (e.g. Scaleway Object Storage, Backblaze B2) for off-site backup.Restore from a snapshot
curl -s -X PUT https://vectors.mydomain.com/collections/documents/snapshots/upload \ -H "api-key: YOUR_KEY" \ -H 'Content-Type: multipart/form-data' \ -F [email protected]The collection is recreated from the snapshot with all its data and original HNSW configuration.
Monitoring with Prometheus and Grafana
Qdrant exposes a /metrics endpoint in Prometheus format on port 6333. Useful production metrics: app_requests_total (throughput), app_requests_duration_seconds (p50/p99 latency), collection_vector_count (corpus size) and gRPC errors.
Add the target to your prometheus.yml:
scrape_configs:
- job_name: qdrant
static_configs:
- targets: ['localhost:6333']
metrics_path: /metrics
bearer_token: YOUR_KEYIn Grafana, import the community Qdrant dashboard (ID 18278) to visualise p99 latency and error rate in real time.
Common troubleshooting
- Connection refused on port 6333: check the container is running (
docker compose ps) and the port is not blocked byufw—ufw allow from 127.0.0.1 to any port 6333. If the port is bound to0.0.0.0instead of127.0.0.1, fix thecompose.ymland restart. - OOM error / container killed: reduce
ef_construct(e.g. 100 → 64) or enable scalar quantization to lower the index RAM footprint. Enableon_disk: truefor original vectors. Check actual consumption withdocker stats qdrant. - Slow startup or first load: the HNSW index is loaded from disk on first access. Use NVMe SSD (not a spinning disk) and warm up the collection after startup with an empty search request.
- Unauthorized 403: the key passed in the
api-keyheader does not matchQDRANT__SERVICE__API_KEY. Check for trailing spaces or newlines in the environment variable. - Degraded recall after enabling binary quantization: increase
oversamplingto3.0or4.0and enablerescore=Trueto re-rank candidates with original vectors.
Pair Qdrant with Ollama for a fully self-contained RAG stack
Run Ollama on the same VPS for embedding generation (ollama pull mxbai-embed-large) and inference (ollama pull llama3.1:8b). Your RAG pipeline then works end-to-end without any external API call: user question → Ollama embed → Qdrant search → top-k chunks → Ollama generate → grounded answer.
Add LiteLLM in front to expose a unified OpenAI-compatible endpoint to your app — and swap models without changing a line of code. A VPS with 8 GB of RAM comfortably hosts Qdrant + Ollama (quantised 7B model) + LiteLLM for a corpus of a few hundred thousand vectors.