Why self-host Whisper on premise on a VPS
Concrete benefits of self-hosting
- Full confidentiality: your audio files never transit through a third party
- Controlled cost: transcribing 60 minutes costs a few cents of CPU, no per-minute billing
- Private API exposed only behind your reverse proxy, protected by key or JWT
- 24/7 availability with no dependency on OpenAI pricing policies
- Model selection based on your RAM constraints: from 1 GB (tiny) to 10 GB (large-v3)
- Simplified GDPR compliance: processing in your datacenter, not at a US subprocessor
Hardware and software prerequisites
For a CPU deployment (small or base model): 2 vCPUs, 4 GB RAM, 20 GB disk. For the medium model: 8 GB RAM recommended. For large-v3 with GPU: 10 GB VRAM minimum (NVIDIA card with CUDA 12+ drivers). On the software side, only Docker Engine 24+ and Docker Compose v2 are needed — no Python or PyTorch to install on the host.
Choosing your Whisper model: RAM, speed and accuracy comparison
Available Whisper models
Scroll the table
| Model | Parameters | RAM required | English WER | Speed (real-time) |
|---|---|---|---|---|
| tiny | 39 M | ~1 GB | ~7.6% | ~10× |
| base | 74 M | ~1 GB | ~5.0% | ~7× |
| small | 244 M | ~2 GB | ~3.4% | ~4× |
| medium | 769 M | ~5 GB | ~2.9% | ~2× |
| large-v3 | 1,550 M | ~10 GB | ~2.4% | ~1× |
| large-v3-turbo | 809 M | ~6 GB | ~2.5% | ~5× |
The small model is an excellent trade-off for a 4 GB VPS without GPU. large-v3-turbo delivers quality close to large-v3 at 5× the speed, for an 8 GB VPS.
Whisper.cpp: the CPU-only alternative without CUDA
If your VPS has no NVIDIA GPU, Whisper.cpp (github.com/ggerganov/whisper.cpp) is the ideal alternative. It is a pure C/C++ port of the Whisper model that runs on CPU without Python, PyTorch or CUDA. It uses SIMD instructions (AVX2 on x86, NEON on ARM) to accelerate inference. On a 4-vCPU VPS, the small model transcribes noticeably faster thanks to CTranslate2's quantized C++ inference engine. The official Whisper.cpp Docker image weighs only ~150 MB versus ~2 GB for the Python image. It is the recommended choice for entry-level VPS instances (2–4 GB RAM).
Step-by-step Docker deployment (openai-whisper-asr-webservice v1.10.0)
Prepare the server
Install Docker Engine and Docker Compose:
curl -fsSL https://get.docker.com | sh && sudo usermod -aG docker $USER. Block port 9000 from the outside:ufw deny 9000/tcp(the API will be accessible via reverse proxy).Create the working directory
mkdir -p ~/whisper && cd ~/whisperWrite the docker-compose.yml
services: whisper: image: onerahmet/openai-whisper-asr-webservice:latest restart: unless-stopped ports: - "127.0.0.1:9000:9000" environment: - ASR_MODEL=small # tiny | base | small | medium | large-v3 - ASR_ENGINE=faster_whisper - ASR_MODEL_PATH=/data/models volumes: - ./models:/data/models # For GPU: replace image with :latest-gpu and add: # deploy: # resources: # reservations: # devices: # - capabilities: [gpu]Choose the model and start
Replace
ASR_MODEL=smallwith the model matching your RAM (see table above). Then:docker compose up -d. On first start, Whisper automatically downloads the model weights into the./modelsvolume (~250 MB for small, ~3 GB for large-v3).Verify the API is running
curl http://127.0.0.1:9000/docs # → Swagger UI; or test a transcription: curl -F "[email protected]" http://127.0.0.1:9000/asr?encode=true&task=transcribe&language=enConfigure the Nginx reverse proxy
location /whisper/ { proxy_pass http://127.0.0.1:9000/; proxy_set_header Authorization $http_authorization; if ($http_x_api_key != "YOUR_SECRET_KEY") { return 403; } }Protect the API with a secret key
Never expose the Whisper API without authentication (see security section). Generate a key:
openssl rand -hex 32. Add it to your proxy and clients.Test a transcription via the secured API
curl -H "X-Api-Key: YOUR_SECRET_KEY" \ -F "[email protected]" \ https://your-domain.com/whisper/asr?task=transcribe&language=en
Real-world integration: connecting Whisper to n8n or Nextcloud Talk
Once the API is operational, you can integrate it into your existing workflows. In n8n, add an HTTP Request node pointing to https://your-domain.com/whisper/asr with the X-Api-Key header and attach your audio file as multipart/form-data. The returned transcription (JSON with segments and timestamps) can then feed a Notion, Google Sheets or custom database node. For Nextcloud Talk, the talk_recording plugin lets you record calls and automatically send the resulting file to your Whisper endpoint — the transcription comes back into the conversation thread within seconds.
CPU vs GPU benchmarks on typical ServOrbit VPS
On a 4-vCPU / 8 GB RAM VPS (faster-whisper engine, small model): one minute of audio is transcribed in approximately 18 seconds, a 3.3× real-time ratio. On an 8-vCPU / 16 GB RAM VPS with the medium model: ~25 seconds per minute of audio (2.4×). With an NVIDIA RTX 4090 GPU and large-v3: ~4 seconds per minute of audio (15×). For batch processing (podcasts, daily meetings), a CPU VPS with 8 GB and the small model offers the optimal performance-to-cost ratio. For real-time transcription (< 5 s latency), a GPU VPS with large-v3-turbo is required.
Security: never expose the Whisper API without authentication
Essential security rules
- Bind the Docker port to 127.0.0.1 only (
127.0.0.1:9000:9000) — never0.0.0.0:9000 - Protect external access with an API key in the reverse proxy (header
X-Api-Key) - Add an Nginx rate-limit:
limit_req_zone $binary_remote_addr zone=whisper:10m rate=10r/m - Enable TLS on your domain (Let's Encrypt) — never transmit audio in plaintext
- Do not store audio files on disk longer than necessary: configure
--tmpfs /tmpin Docker - Restrict the endpoint to internal IPs if the service is only used internally
Troubleshooting: common errors
Common issues and solutions
- CUDA not found / RuntimeError: CUDA error: use the
:latest(CPU) image instead of:latest-gpu, or verify thatnvidia-container-toolkitis installed and CUDA 12+ drivers are present (nvidia-smi). - Model not downloaded / FileNotFoundError: the
./modelsvolume is not mounted correctly, or the initial download was interrupted. Delete the contents of./modelsand restartdocker compose up -d— Whisper re-downloads automatically. - OOMKilled / Out of memory: the chosen model exceeds available RAM. Switch to a smaller model (
medium→small,small→base) or increase your VPS RAM. - 422 Unprocessable Entity on /asr: the audio file must be sent as
multipart/form-datawith the field namedaudio_file. Also check thatencode=trueis present in the query string for MP3/MP4 formats.
For VPS instances without a GPU, combine Whisper.cpp (CPU-only, ~150 MB image) with the 4-bit quantized small model (-m models/ggml-small-q5_1.bin): you cut RAM usage by ~2.5× while retaining accuracy close to the float16 model.
Deploy Whisper from the ServOrbit Marketplace
ServOrbit offers a pre-configured Whisper application in its AI Marketplace. One click and your VPS is ready: Docker installed, onerahmet/openai-whisper-asr-webservice:latest deployed, Nginx configured with TLS and an automatically generated API key. You choose the model (tiny to large-v3) at install time and benefit from one-click updates.