Deployment guide

Hosting Qdrant on a VPS

Deploy on a VPS Cloud →

Tutorial

Hosting Qdrant on a VPS

Databases8 min read15 steps

Qdrant is a vector database written in Rust, built for semantic search, recommendation systems and the RAG pipelines of AI applications. Lightweight and fast, it self-hosts perfectly on a VPS to keep your embeddings under your control and avoid the per-vector billing of managed services. This guide covers the full path: installation with Docker Compose, securing with an API key and TLS reverse proxy, inserting vectors in Python with the official v1 client, tuning the HNSW index and quantization, backing up with snapshots and monitoring via the Prometheus-compatible `/metrics` endpoint. ServOrbit NVMe SSD VPS plans start at €3.99/month and handle corpora under one million vectors without any special configuration.

Contents· Why self-host Qdrant on a VPS1/14
  1. 01Why self-host Qdrant on a VPS
  2. 02Prerequisites and recommended VPS resources
  3. 03Install Qdrant with Docker Compose
  4. 04Securing access: API key and TLS
  5. 05Test authentication and the reverse proxy
  6. 06Create a collection and insert vectors with the Python v1 client
  7. 07Optimise performance: HNSW tuning
  8. 08Key HNSW parameters
  9. 09Configure HNSW at collection creation
  10. 10Reduce memory footprint: scalar and binary quantization
  11. 11Back up and restore with Qdrant snapshots
  12. 12Monitoring with Prometheus and Grafana
  13. 13Common troubleshooting
  14. 14Pair Qdrant with Ollama for a fully self-contained RAG stack

Why self-host Qdrant on a VPS

Qdrant stores and queries embedding vectors to find the most semantically similar items: it is the search building block at the heart of RAG chatbots, recommendation engines and natural-language search. Written in Rust, it is fast and resource-efficient, making it ideal to host on a VPS rather than paying for a managed vector service billed per stored vector.

Self-hosting is essential for AI applications handling sensitive data: your embeddings, often derived from confidential internal documents, never leave your infrastructure. You also control latency, which is decisive when every user request triggers a vector search, and you avoid the rate limits of third-party APIs during load spikes.

Prerequisites and recommended VPS resources

  • RAM: minimum 1 GB for testing, 2 GB for a few hundred thousand vectors, 4–8 GB beyond a million high-dimensional vectors (the HNSW index lives in memory).
  • Storage: 20–50 GB of NVMe SSD to persist collections and snapshots.
  • CPU: 1–2 vCPUs are sufficient for moderate traffic; add cores if you index continuously.
  • Docker and Docker Compose installed on the VPS.
  • Ports 6333 (REST) and 6334 (gRPC) accessible from your application; 80/443 for the public reverse proxy.
  • A DNS A record (e.g. vectors.mydomain.com) if you expose the API over HTTPS.

Install Qdrant with Docker Compose

  1. Create the directory and compose.yml

    On the VPS, create a dedicated folder: mkdir -p /opt/qdrant && cd /opt/qdrant. Create the following compose.yml:

    services:
      qdrant:
        image: qdrant/qdrant:latest
        restart: unless-stopped
        ports:
          - "127.0.0.1:6333:6333"
          - "127.0.0.1:6334:6334"
        volumes:
          - ./qdrant_storage:/qdrant/storage
        environment:
          - QDRANT__SERVICE__API_KEY=${QDRANT_API_KEY}

    Binding ports to 127.0.0.1 prevents any direct exposure on the Internet; only the reverse proxy can reach them.

  2. Generate the API key and start the container

    Create the .env file with your key:

    echo "QDRANT_API_KEY=$(openssl rand -hex 32)" > .env

    Start Qdrant:

    docker compose up -d
  3. Check instance health

    Wait a few seconds then verify the health endpoint:

    curl -s http://localhost:6333/healthz
    # Expected: {"title":"qdrant - vector search engine","version":"..."}
  4. Create your first collection

    Create a test collection with 384-dimensional vectors (typical format for all-MiniLM-L6-v2):

    curl -s -X PUT http://localhost:6333/collections/test \
      -H 'Content-Type: application/json' \
      -H "api-key: $(grep QDRANT_API_KEY .env | cut -d= -f2)" \
      -d '{"vectors":{"size":384,"distance":"Cosine"}}'

Securing access: API key and TLS

Qdrant enables no authentication by default: an instance exposed without an API key would allow anyone to read, modify or delete all your collections. The QDRANT__SERVICE__API_KEY environment variable is the first line of defence; the TLS reverse proxy is the second.

For Caddy, the minimal configuration is:

vectors.mydomain.com {
    reverse_proxy localhost:6333
}

Caddy automatically obtains and renews the Let's Encrypt certificate. For nginx, add inside your server block:

location / {
    proxy_pass http://127.0.0.1:6333;
    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
}

Enable TLS with Certbot: certbot --nginx -d vectors.mydomain.com.

Test authentication and the reverse proxy

  1. Verify that unauthenticated access is rejected

    curl -s https://vectors.mydomain.com/collections
    # Expected: {"status":{"error":"Unauthorized",...}}
  2. Authenticated call with the API key

    curl -s https://vectors.mydomain.com/collections \
      -H "api-key: YOUR_KEY"
    # Expected: {"result":{"collections":[]},"status":"ok",...}
  3. Verify the TLS certificate

    curl -sv https://vectors.mydomain.com/healthz 2>&1 | grep 'SSL connection\|issuer'
    # Expected: lines confirming Let's Encrypt and TLS 1.3

Create a collection and insert vectors with the Python v1 client

  1. Install the official client

    pip install qdrant-client>=1.7
  2. Connect with the API key

    from qdrant_client import QdrantClient
    
    client = QdrantClient(
        url="https://vectors.mydomain.com",
        api_key="YOUR_KEY",
    )
  3. Create a collection with VectorParams

    from qdrant_client.models import Distance, VectorParams
    
    client.recreate_collection(
        collection_name="documents",
        vectors_config=VectorParams(size=384, distance=Distance.COSINE),
    )

    size must match the embedding dimension of your model. Distance.COSINE suits most sentence-transformer models.

  4. Insert points with payload and query

    from qdrant_client.models import PointStruct
    
    # Insert vectors
    client.upsert(
        collection_name="documents",
        points=[
            PointStruct(
                id=1,
                vector=[0.1, 0.2, ...],  # vector of size 384
                payload={"title": "Qdrant Guide", "lang": "en"},
            ),
        ],
    )
    
    # Search for the 5 nearest neighbours
    results = client.search(
        collection_name="documents",
        query_vector=[0.1, 0.21, ...],
        limit=5,
        with_payload=True,
    )
    for r in results:
        print(r.score, r.payload)

Optimise performance: HNSW tuning

The HNSW (Hierarchical Navigable Small World) algorithm is the core of fast search in Qdrant. Three parameters govern the speed/accuracy/memory trade-off.

Key HNSW parameters

Scroll the table

ParameterDefaultRecommended rangeImpact
m1616–64Connections per node — higher = better recall, more RAM
ef_construct100100–400Index build quality — higher = better index, slower build
ef128 (at query)64–512Search quality — higher = better recall, slower query

Configure HNSW at collection creation

Pass HNSW parameters directly in recreate_collection:

from qdrant_client.models import HnswConfigDiff

client.recreate_collection(
    collection_name="documents",
    vectors_config=VectorParams(size=384, distance=Distance.COSINE),
    hnsw_config=HnswConfigDiff(m=32, ef_construct=200),
)

For balanced search on a corpus of 500,000 vectors, m=32 and ef_construct=200 provide good recall (>0.97) without overloading RAM. Increase ef at query time for cases where precision matters more than latency.

Reduce memory footprint: scalar and binary quantization

  • Scalar quantization (uint8): compresses each float32 coordinate (4 bytes) to uint8 (1 byte), roughly a 4× RAM reduction. Slight precision loss offset by oversampling. Enable with ScalarQuantizationConfig(type=ScalarType.INT8, quantile=0.99, always_ram=True).
  • Binary quantization: up to 32× reduction by binarising each dimension. Ideal for models trained for binary quantization (e.g. Cohere embed-english-v3.0). Faster but lower recall without oversampling.
  • on_disk parameter: store the original vectors on disk (on_disk=True) and keep only the quantised index in RAM. Essential on memory-constrained VPS.
  • Oversampling: with oversampling=2.0 at query time, Qdrant retrieves 2× more candidates from the quantised index then re-ranks them with the original vectors — recall close to full-precision at moderate cost.

Back up and restore with Qdrant snapshots

  1. Create a snapshot of a collection

    curl -s -X POST https://vectors.mydomain.com/collections/documents/snapshots \
      -H "api-key: YOUR_KEY"
    # Returns: {"result":{"name":"documents-...","creation_time":"..."}}
  2. List available snapshots

    curl -s https://vectors.mydomain.com/collections/documents/snapshots \
      -H "api-key: YOUR_KEY"
  3. Download the snapshot

    curl -O https://vectors.mydomain.com/collections/documents/snapshots/SNAPSHOT_NAME \
      -H "api-key: YOUR_KEY"

    Then copy the .snapshot file to object storage (e.g. Scaleway Object Storage, Backblaze B2) for off-site backup.

  4. Restore from a snapshot

    curl -s -X PUT https://vectors.mydomain.com/collections/documents/snapshots/upload \
      -H "api-key: YOUR_KEY" \
      -H 'Content-Type: multipart/form-data' \
      -F [email protected]

    The collection is recreated from the snapshot with all its data and original HNSW configuration.

Monitoring with Prometheus and Grafana

Qdrant exposes a /metrics endpoint in Prometheus format on port 6333. Useful production metrics: app_requests_total (throughput), app_requests_duration_seconds (p50/p99 latency), collection_vector_count (corpus size) and gRPC errors.

Add the target to your prometheus.yml:

scrape_configs:
  - job_name: qdrant
    static_configs:
      - targets: ['localhost:6333']
    metrics_path: /metrics
    bearer_token: YOUR_KEY

In Grafana, import the community Qdrant dashboard (ID 18278) to visualise p99 latency and error rate in real time.

Common troubleshooting

  • Connection refused on port 6333: check the container is running (docker compose ps) and the port is not blocked by ufw — ufw allow from 127.0.0.1 to any port 6333. If the port is bound to 0.0.0.0 instead of 127.0.0.1, fix the compose.yml and restart.
  • OOM error / container killed: reduce ef_construct (e.g. 100 → 64) or enable scalar quantization to lower the index RAM footprint. Enable on_disk: true for original vectors. Check actual consumption with docker stats qdrant.
  • Slow startup or first load: the HNSW index is loaded from disk on first access. Use NVMe SSD (not a spinning disk) and warm up the collection after startup with an empty search request.
  • Unauthorized 403: the key passed in the api-key header does not match QDRANT__SERVICE__API_KEY. Check for trailing spaces or newlines in the environment variable.
  • Degraded recall after enabling binary quantization: increase oversampling to 3.0 or 4.0 and enable rescore=True to re-rank candidates with original vectors.

Pair Qdrant with Ollama for a fully self-contained RAG stack

Run Ollama on the same VPS for embedding generation (ollama pull mxbai-embed-large) and inference (ollama pull llama3.1:8b). Your RAG pipeline then works end-to-end without any external API call: user question → Ollama embed → Qdrant search → top-k chunks → Ollama generate → grounded answer.

Add LiteLLM in front to expose a unified OpenAI-compatible endpoint to your app — and swap models without changing a line of code. A VPS with 8 GB of RAM comfortably hosts Qdrant + Ollama (quantised 7B model) + LiteLLM for a corpus of a few hundred thousand vectors.

Launch your Qdrant vector database for AI

The ServOrbit Cloud VPS offers the memory and SSD disks you need to host Qdrant and power your RAG pipelines and semantic search engines, with SSL and authentication under your control.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab