Deployment guide

Whisper on premise on a VPS: complete 2026 guide

Deploy on a VPS Cloud →

Tutorial

Whisper on premise on a VPS: complete 2026 guide

Artificial Intelligence6 min read8 steps

Whisper is OpenAI's open-source audio transcription model. Hosting it on premise on your VPS ensures your recordings never leave your infrastructure: full confidentiality, controlled latency and predictable cost. This guide covers step-by-step Docker deployment, model selection based on your RAM, the CPU-only Whisper.cpp alternative, integration with n8n, benchmarks on typical VPS instances, and essential security measures.

Contents· Why self-host Whisper on premise on a VPS1/14
  1. 01Why self-host Whisper on premise on a VPS
  2. 02Concrete benefits of self-hosting
  3. 03Hardware and software prerequisites
  4. 04Choosing your Whisper model: RAM, speed and accuracy comparison
  5. 05Available Whisper models
  6. 06Whisper.cpp: the CPU-only alternative without CUDA
  7. 07Step-by-step Docker deployment (openai-whisper-asr-webservice v1.10.0)
  8. 08Real-world integration: connecting Whisper to n8n or Nextcloud Talk
  9. 09CPU vs GPU benchmarks on typical ServOrbit VPS
  10. 10Security: never expose the Whisper API without authentication
  11. 11Essential security rules
  12. 12Troubleshooting: common errors
  13. 13Common issues and solutions
  14. 14Deploy Whisper from the ServOrbit Marketplace

Why self-host Whisper on premise on a VPS

Concrete benefits of self-hosting

  • Full confidentiality: your audio files never transit through a third party
  • Controlled cost: transcribing 60 minutes costs a few cents of CPU, no per-minute billing
  • Private API exposed only behind your reverse proxy, protected by key or JWT
  • 24/7 availability with no dependency on OpenAI pricing policies
  • Model selection based on your RAM constraints: from 1 GB (tiny) to 10 GB (large-v3)
  • Simplified GDPR compliance: processing in your datacenter, not at a US subprocessor

Hardware and software prerequisites

For a CPU deployment (small or base model): 2 vCPUs, 4 GB RAM, 20 GB disk. For the medium model: 8 GB RAM recommended. For large-v3 with GPU: 10 GB VRAM minimum (NVIDIA card with CUDA 12+ drivers). On the software side, only Docker Engine 24+ and Docker Compose v2 are needed — no Python or PyTorch to install on the host.

Choosing your Whisper model: RAM, speed and accuracy comparison

Available Whisper models

Scroll the table

ModelParametersRAM requiredEnglish WERSpeed (real-time)
tiny39 M~1 GB~7.6%~10×
base74 M~1 GB~5.0%~7×
small244 M~2 GB~3.4%~4×
medium769 M~5 GB~2.9%~2×
large-v31,550 M~10 GB~2.4%~1×
large-v3-turbo809 M~6 GB~2.5%~5×

The small model is an excellent trade-off for a 4 GB VPS without GPU. large-v3-turbo delivers quality close to large-v3 at 5× the speed, for an 8 GB VPS.

Whisper.cpp: the CPU-only alternative without CUDA

If your VPS has no NVIDIA GPU, Whisper.cpp (github.com/ggerganov/whisper.cpp) is the ideal alternative. It is a pure C/C++ port of the Whisper model that runs on CPU without Python, PyTorch or CUDA. It uses SIMD instructions (AVX2 on x86, NEON on ARM) to accelerate inference. On a 4-vCPU VPS, the small model transcribes noticeably faster thanks to CTranslate2's quantized C++ inference engine. The official Whisper.cpp Docker image weighs only ~150 MB versus ~2 GB for the Python image. It is the recommended choice for entry-level VPS instances (2–4 GB RAM).

Step-by-step Docker deployment (openai-whisper-asr-webservice v1.10.0)

  1. Prepare the server

    Install Docker Engine and Docker Compose: curl -fsSL https://get.docker.com | sh && sudo usermod -aG docker $USER. Block port 9000 from the outside: ufw deny 9000/tcp (the API will be accessible via reverse proxy).

  2. Create the working directory

    mkdir -p ~/whisper && cd ~/whisper
  3. Write the docker-compose.yml

    services:
      whisper:
        image: onerahmet/openai-whisper-asr-webservice:latest
        restart: unless-stopped
        ports:
          - "127.0.0.1:9000:9000"
        environment:
          - ASR_MODEL=small          # tiny | base | small | medium | large-v3
          - ASR_ENGINE=faster_whisper
          - ASR_MODEL_PATH=/data/models
        volumes:
          - ./models:/data/models
        # For GPU: replace image with :latest-gpu and add:
        # deploy:
        #   resources:
        #     reservations:
        #       devices:
        #         - capabilities: [gpu]
  4. Choose the model and start

    Replace ASR_MODEL=small with the model matching your RAM (see table above). Then: docker compose up -d. On first start, Whisper automatically downloads the model weights into the ./models volume (~250 MB for small, ~3 GB for large-v3).

  5. Verify the API is running

    curl http://127.0.0.1:9000/docs
    # → Swagger UI; or test a transcription:
    curl -F "[email protected]" http://127.0.0.1:9000/asr?encode=true&task=transcribe&language=en
  6. Configure the Nginx reverse proxy

    location /whisper/ {
        proxy_pass http://127.0.0.1:9000/;
        proxy_set_header Authorization $http_authorization;
        if ($http_x_api_key != "YOUR_SECRET_KEY") {
            return 403;
        }
    }
  7. Protect the API with a secret key

    Never expose the Whisper API without authentication (see security section). Generate a key: openssl rand -hex 32. Add it to your proxy and clients.

  8. Test a transcription via the secured API

    curl -H "X-Api-Key: YOUR_SECRET_KEY" \
      -F "[email protected]" \
      https://your-domain.com/whisper/asr?task=transcribe&language=en

Real-world integration: connecting Whisper to n8n or Nextcloud Talk

Once the API is operational, you can integrate it into your existing workflows. In n8n, add an HTTP Request node pointing to https://your-domain.com/whisper/asr with the X-Api-Key header and attach your audio file as multipart/form-data. The returned transcription (JSON with segments and timestamps) can then feed a Notion, Google Sheets or custom database node. For Nextcloud Talk, the talk_recording plugin lets you record calls and automatically send the resulting file to your Whisper endpoint — the transcription comes back into the conversation thread within seconds.

CPU vs GPU benchmarks on typical ServOrbit VPS

On a 4-vCPU / 8 GB RAM VPS (faster-whisper engine, small model): one minute of audio is transcribed in approximately 18 seconds, a 3.3× real-time ratio. On an 8-vCPU / 16 GB RAM VPS with the medium model: ~25 seconds per minute of audio (2.4×). With an NVIDIA RTX 4090 GPU and large-v3: ~4 seconds per minute of audio (15×). For batch processing (podcasts, daily meetings), a CPU VPS with 8 GB and the small model offers the optimal performance-to-cost ratio. For real-time transcription (< 5 s latency), a GPU VPS with large-v3-turbo is required.

Security: never expose the Whisper API without authentication

Essential security rules

  • Bind the Docker port to 127.0.0.1 only (127.0.0.1:9000:9000) — never 0.0.0.0:9000
  • Protect external access with an API key in the reverse proxy (header X-Api-Key)
  • Add an Nginx rate-limit: limit_req_zone $binary_remote_addr zone=whisper:10m rate=10r/m
  • Enable TLS on your domain (Let's Encrypt) — never transmit audio in plaintext
  • Do not store audio files on disk longer than necessary: configure --tmpfs /tmp in Docker
  • Restrict the endpoint to internal IPs if the service is only used internally

Troubleshooting: common errors

Common issues and solutions

  • CUDA not found / RuntimeError: CUDA error: use the :latest (CPU) image instead of :latest-gpu, or verify that nvidia-container-toolkit is installed and CUDA 12+ drivers are present (nvidia-smi).
  • Model not downloaded / FileNotFoundError: the ./models volume is not mounted correctly, or the initial download was interrupted. Delete the contents of ./models and restart docker compose up -d — Whisper re-downloads automatically.
  • OOMKilled / Out of memory: the chosen model exceeds available RAM. Switch to a smaller model (medium → small, small → base) or increase your VPS RAM.
  • 422 Unprocessable Entity on /asr: the audio file must be sent as multipart/form-data with the field named audio_file. Also check that encode=true is present in the query string for MP3/MP4 formats.

For VPS instances without a GPU, combine Whisper.cpp (CPU-only, ~150 MB image) with the 4-bit quantized small model (-m models/ggml-small-q5_1.bin): you cut RAM usage by ~2.5× while retaining accuracy close to the float16 model.

Deploy Whisper from the ServOrbit Marketplace

ServOrbit offers a pre-configured Whisper application in its AI Marketplace. One click and your VPS is ready: Docker installed, onerahmet/openai-whisper-asr-webservice:latest deployed, Nginx configured with TLS and an automatically generated API key. You choose the model (tiny to large-v3) at install time and benefit from one-click updates.

Deploy Whisper on a ServOrbit VPS in 1 click

Whisper is available on the ServOrbit Marketplace: Docker, ffmpeg and the network configuration are managed automatically. Your transcriptions stay on your VPS — no data is sent to a third-party service.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab