Why self-host ComfyUI on a VPS
ComfyUI organizes image generation into node graphs: each step (model loading, prompt encoding, sampling, VAE) is a linkable, reusable block, which makes workflows reproducible and shareable in JSON format. Unlike an online generation service, self-hosting gives you control over the checkpoint models, the LoRAs, the ControlNets and the extensions, without censorship or quota. On a GPU VPS, you get a studio available 24/7 that an entire creative team can use remotely, and whose API lets you industrialize generation from your own scripts or pipelines.
Concrete benefits of self-hosting
- Reproducible node-based workflows, exportable as JSON and shareable within the team
- Free library of checkpoints, LoRAs and ControlNets with no quota or censorship
- HTTP API to automate generation from your scripts and pipelines
- Remote GPU accessible 24/7 without tying up a local workstation
- Installation of custom nodes (community extensions) without restriction
- Cost control: a GPU VPS by the hour or the month rather than paying per image
Hardware requirements by use case
Sizing depends directly on the execution mode and target models. In CPU mode (slow, for testing and prototyping), a 4 vCPU VPS with 8 GB RAM is enough to load an SDXL checkpoint, but expect several minutes per image. In GPU mode, the bottleneck is VRAM: 8 GB of NVIDIA VRAM can run SDXL in fp16 with the --lowvram flag, which offloads text encoders to system RAM; 12 to 16 GB of VRAM are recommended for Flux.1 at full precision. For storage: an SDXL checkpoint weighs around 6 to 7 GB, Flux.1 schnell (full precision) is 23.8 GB, its fp8 version 17.2 GB. Plan for at least 50 GB of SSD, ideally 100 GB if you intend to store multiple models and their LoRAs. System RAM must be at least 16 GB when GPU-to-RAM offloading is active.
Minimum configuration by scenario
Scroll the table
| Scenario | vCPU | RAM | VRAM | Disk |
|---|---|---|---|---|
| CPU test (SDXL, slow) | 4 | 8 GB | — (no GPU) | 50 GB |
| GPU SDXL comfortable | 4 | 16 GB | 8 GB NVIDIA | 80 GB |
| GPU Flux.1 (recommended) | 8 | 32 GB | 16 GB NVIDIA | 100 GB |
| Production multi-user | 8+ | 32 GB+ | 24 GB NVIDIA | 200 GB+ |
Two installation methods: Docker vs Python venv
ComfyUI can be deployed two ways: via Docker (isolation, reproducibility, simplified GPU dependency management) or via a Python virtual environment (closer to bare-metal, more flexible for experimental custom nodes). On a production VPS, Docker is recommended for ease of maintenance and version isolation.
Method A — Python venv installation (direct hardware access)
Install system dependencies
On Ubuntu 22.04/24.04:
apt update && apt install -y git python3.12 python3.12-venv python3-pip. ComfyUI supports Python 3.12 and 3.13; version 3.13 is very well supported, 3.14 may cause compatibility issues with some custom nodes.Clone the repository and create the venv
git clone https://github.com/comfyanonymous/ComfyUI.git /opt/comfyui && cd /opt/comfyui && python3.12 -m venv venv && source venv/bin/activate && pip install -r requirements.txtInstall PyTorch with CUDA or CPU support
For NVIDIA GPU (CUDA):
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu124. For CPU-only mode:pip install torch torchvision. PyTorch 2.7 is the minimum supported version; a newer version is strongly recommended.Launch ComfyUI
GPU mode:
python main.py --listen 0.0.0.0. CPU mode:python main.py --cpu --listen 0.0.0.0. The--listen 0.0.0.0flag exposes ComfyUI on all network interfaces of the VPS (required for access via tunnel or reverse proxy). The interface is available on port 8188.
Method B — Docker deployment with GPU
Prepare the GPU VPS
On a VPS with an NVIDIA GPU, install the drivers then the NVIDIA Container Toolkit so that Docker can access the GPU. Validate with
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi.Run ComfyUI in a container
Start a ComfyUI image with GPU access and persistent volumes:
docker run -d --gpus all -p 127.0.0.1:8188:8188 -v /opt/comfyui/models:/app/models -v /opt/comfyui/output:/app/output --name comfyui ghcr.io/ai-dock/comfyui:latest-cuda. Restricting to 127.0.0.1 avoids direct exposure. Adjust the image tag to match your VPS CUDA version.Verify the GPU is detected
After startup:
docker logs comfyui | grep -i 'cuda\|gpu\|device'. ComfyUI displays the selected device at startup. If you seeUsing CPU, your GPU is not accessible from the container — check the NVIDIA Container Toolkit.
Downloading models from HuggingFace
Models are downloaded from HuggingFace using wget or the HuggingFace CLI (pip install huggingface_hub). Each file type has its dedicated folder in the ComfyUI tree. For SDXL: place the checkpoint .safetensors file in models/checkpoints/. For Flux.1: the architecture differs — the diffusion model goes in models/diffusion_models/ (or models/unet/ depending on the version), and Flux requires two text encoders in models/text_encoders/: clip_l.safetensors and t5xxl_fp16.safetensors (or t5xxl_fp8_e4m3fn_scaled.safetensors to save VRAM). The VAE (ae.safetensors) goes in models/vae/. Flux.1 schnell is freely available from black-forest-labs/FLUX.1-schnell on HuggingFace (23.8 GB full precision, 17.2 GB fp8). Flux.1 dev is gated — you must accept the terms of use on HuggingFace before downloading.
Nginx reverse proxy with authentication
Create the basic auth file
apt install -y apache2-utils && htpasswd -c /etc/nginx/.htpasswd your_user. ComfyUI has no native authentication: without this step, your instance is open to everyone.Configure the Nginx virtual host
Create
/etc/nginx/sites-available/comfyuiwith:server { listen 443 ssl; server_name comfy.yourdomain.com; ssl_certificate /etc/letsencrypt/live/comfy.yourdomain.com/fullchain.pem; ssl_certificate_key /etc/letsencrypt/live/comfy.yourdomain.com/privkey.pem; auth_basic "ComfyUI"; auth_basic_user_file /etc/nginx/.htpasswd; location / { proxy_pass http://127.0.0.1:8188; proxy_read_timeout 300s; proxy_send_timeout 300s; proxy_http_version 1.1; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; proxy_set_header Host $host; } }. The WebSocket upgrade is mandatory for ComfyUI's real-time API.Obtain the Let's Encrypt certificate and enable
certbot --nginx -d comfy.yourdomain.com && ln -s /etc/nginx/sites-available/comfyui /etc/nginx/sites-enabled/ && nginx -t && systemctl reload nginx. Then close port 8188 at the firewall:ufw deny 8188.
Installing ComfyUI Manager and custom nodes
ComfyUI Manager is the essential extension for managing custom nodes from the graphical interface. In Python venv: cd /opt/comfyui/custom_nodes && git clone https://github.com/Comfy-Org/ComfyUI-Manager.git && cd ComfyUI-Manager && pip install -r requirements.txt. Then relaunch ComfyUI with python main.py --enable-manager --listen 0.0.0.0. A "Manager" icon appears in the interface: you can install, update and disable popular custom nodes (WAS Node Suite, ControlNet Preprocessors, IP-Adapter, etc.) without any command line. In Docker, mount a volume on custom_nodes/ so that installations survive container restarts.
ComfyUI versus AUTOMATIC1111 (Stable Diffusion WebUI)
Scroll the table
| Criterion | ComfyUI | AUTOMATIC1111 |
|---|---|---|
| Approach | Visual node-based workflows | Classic tabbed interface |
| Reproducibility | Excellent (workflow exported as JSON) | Limited to the entered parameters |
| VRAM consumption | Optimized, handles small GPUs better | More demanding at equal configuration |
| Learning curve | Steeper (graph logic) | More approachable for beginners |
| API automation | Native and granular | API present but less flexible |
| Recent models (Flux, SD3) | Fast, reference-grade support | Often later support |
| Custom nodes / extensions | Very rich node ecosystem | Large extension catalog |
| Ideal use case | Advanced pipelines and automation | Fast interactive generation |
Troubleshooting: 4 common errors
On a freshly configured VPS, several errors come up consistently. Here are the causes and fixes.
Frequent errors and solutions
- CUDA not available / Using CPU: ComfyUI did not detect a GPU. Causes: PyTorch installed without CUDA support (
pip install torchwithout the CUDA index), or missing NVIDIA drivers. Check withpython -c "import torch; print(torch.cuda.is_available())". IfFalse, reinstall PyTorch with--index-url https://download.pytorch.org/whl/cu124. In Docker, verify that the NVIDIA Container Toolkit is installed and that you launch with--gpus all. - CUDA out of memory (OOM): the model does not fit in VRAM. Add
--lowvramwhen launching ComfyUI: this flag forces text encoders to be offloaded to system RAM. For Flux on 8 GB VRAM, also use the fp8 variant of the model. As a last resort,--novramoffloads everything to RAM (very slow). Reducing the generation resolution (512×512 instead of 1024×1024) also helps immediately. - ERROR: Could not find model / model not found: the file is not in the right place. ComfyUI looks for checkpoints in
models/checkpoints/, Flux diffusion models inmodels/diffusion_models/(ormodels/unet/), text encoders inmodels/text_encoders/. A.safetensorsfile in the wrong subfolder will not appear in the interface. Refresh the list with the "Refresh" button in the model loading node. - Port 8188 already in use: a ComfyUI process or another application is already using the port.
lsof -i :8188identifies the PID. Launch ComfyUI on another port with--port 8189and update your Nginx config accordingly. In Docker, the conflict may come from a stopped but not removed container:docker rm comfyuibefore restarting.
To industrialize generation, leverage the API: submit your workflows via POST to /prompt and retrieve the results through the /ws WebSocket that notifies the end of each task. The --lowvram flag at launch allows SDXL models to run on 8 GB VRAM: ComfyUI intelligently offloads text encoders to system RAM. For remote access without a certificate (development), use an SSH tunnel: ssh -L 8188:localhost:8188 user@your-vps — ComfyUI stays accessible on http://localhost:8188 from your workstation without any public exposure.