Deployment guide

VPS Monitoring with Grafana and Prometheus

Deploy on a VPS Cloud →

Tutorial

VPS Monitoring with Grafana and Prometheus

Security & Monitoring11 min read8 steps

When you want historical metrics, custom dashboards and fine-grained alerting, the Prometheus + Grafana duo is the open source reference. Prometheus 3.x collects and stores time series, Grafana 13.x visualizes them. Here is how to deploy this complete stack — Node Exporter, cAdvisor, Alertmanager included — on your own VPS.

Contents· Why self-host Prometheus and Grafana on a VPS1/11
  1. 01Why self-host Prometheus and Grafana on a VPS
  2. 02What this stack brings to your VPS
  3. 03VPS prerequisites and resources to plan for
  4. 04Deploying the stack with Docker Compose
  5. 05Importing community Grafana dashboards
  6. 06Configuring alerts with Alertmanager
  7. 07Prometheus vs InfluxDB vs Zabbix: which to choose?
  8. 08Hardening: never expose Prometheus publicly
  9. 09Adding Docker container monitoring with cAdvisor
  10. 10Troubleshooting: the most common errors
  11. 11Going further

Why self-host Prometheus and Grafana on a VPS

Unlike an all-in-one tool, the Prometheus + Grafana stack clearly separates collection (Prometheus scrapes metrics exposed by exporters), storage (built-in time series database) and visualization (Grafana). This modularity is precisely what makes it a standard: you monitor the VPS itself with Node Exporter, your containers with cAdvisor, your PostgreSQL or MySQL database with the dedicated exporter, and you instrument your own applications in the Prometheus format.

All of this on a VPS you control, which spares you from sending sensitive infrastructure metrics to a SaaS billed by ingestion. You keep the history for as long as your disk allows, you build dashboards specific to your stack, and you define your own alerting rules via Alertmanager.

With the release of Prometheus 3.x (latest stable: v3.15.0, September 2026), the TSDB storage engine has been optimized to reduce memory consumption by 15 to 20% on modest machine fleets — a direct advantage for VPS. Grafana 13.x (v13.2.3, late September 2026) brings a redesigned dashboard composer and a unified alerting system. This is the approach to favor when you go beyond simply 'is the site responding?' to move into fine-grained analysis of load trends.

What this stack brings to your VPS

  • Detailed system metrics (CPU, RAM, disk I/O, network) via Node Exporter v1.12.1, kept as long-term history.
  • Monitoring of Docker containers with cAdvisor: per-container consumption in real time.
  • Custom Grafana dashboards, importable from a community library of thousands of templates.
  • Powerful alerting via Alertmanager v0.34.1: thresholds, grouping, silences and multi-channel routing (email, Slack, PagerDuty...).
  • HTTP/TCP endpoint monitoring with Blackbox Exporter: response time detection and TLS certificate expiry.
  • Application log collection with Loki (same vendor as Grafana): correlate logs and metrics in a single dashboard.
  • PromQL query language to create derived indicators and rates of change.
  • Instrumentation of your own applications in the Prometheus format, with no external ingestion cost.

VPS prerequisites and resources to plan for

This stack is more demanding than a simple monitor because Prometheus keeps time series in memory before writing them to disk. Plan for a VPS with 2 vCPU and 2 to 4 GB of RAM to monitor one to several servers. Add Loki and Alertmanager and aim for 4 GB minimum.

Storage is the critical point: Prometheus generates approximately 1 to 2 MB per time series per day depending on scrape frequency (15 s by default). For 500 series over 30 days, plan for 15 to 30 GB. Start with 40 GB of SSD, and adjust --storage.tsdb.retention.time according to your needs.

Software prerequisites: Docker 26+ and Docker Compose v2, a domain name pointing to your VPS (for the HTTPS reverse proxy for Grafana), and ufw or equivalent to restrict ports.

Deploying the stack with Docker Compose

  1. Create the project directory structure

    On your VPS, create a dedicated folder and the file structure:

    mkdir -p ~/monitoring/{prometheus,alertmanager,loki}
    cd ~/monitoring

    This organization isolates each configuration in its own directory, simplifying updates.

  2. Write the docker-compose.yml file

    Define the following services in a docker-compose.yml file: prometheus (image prom/prometheus:v3.15.0), grafana (image grafana/grafana:13.2.3), node-exporter (image prom/node-exporter:v1.12.1), cadvisor (image gcr.io/cadvisor/cadvisor:latest), alertmanager (image prom/alertmanager:v0.34.1).

    Mount configuration files as bind-mount volumes and create named volumes to persist data: prometheus_data, grafana_data. Example bind for node-exporter:

    volumes:
      - /proc:/host/proc:ro
      - /sys:/host/sys:ro
      - /:/rootfs:ro

    Then run docker compose up -d to start the stack.

  3. Configure prometheus.yml with all scrape jobs

    In prometheus/prometheus.yml, declare targets in the scrape_configs section. Four essential jobs:

    scrape_configs:
      - job_name: 'node'
        static_configs:
          - targets: ['node-exporter:9100']
      - job_name: 'cadvisor'
        static_configs:
          - targets: ['cadvisor:8080']
      - job_name: 'alertmanager'
        static_configs:
          - targets: ['alertmanager:9093']
      - job_name: 'prometheus'
        static_configs:
          - targets: ['localhost:9090']

    Reload configuration without restarting: curl -X POST http://localhost:9090/-/reload (requires --web.enable-lifecycle flag).

  4. Configure Alertmanager (alertmanager.yml)

    Create alertmanager/alertmanager.yml. A minimal configuration routing alerts to Slack:

    route:
      receiver: 'slack-notifications'
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h
    receivers:
      - name: 'slack-notifications'
        slack_configs:
          - api_url: 'https://hooks.slack.com/services/YOUR/WEBHOOK'
            channel: '#vps-alerts'
            title: '{{ .CommonAnnotations.summary }}'

    For email, replace slack_configs with email_configs with to, from, smarthost, auth_username and auth_password fields.

  5. Create Prometheus alerting rules

    In prometheus/rules.yml, define your rules. Three essentials for a VPS:

    groups:
      - name: vps
        rules:
          - alert: HighCPU
            expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
            for: 5m
            annotations:
              summary: 'CPU > 85% for 5 min'
          - alert: LowMemory
            expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
            for: 5m
            annotations:
              summary: 'Available memory < 10%'
          - alert: DiskFull
            expr: (node_filesystem_size_bytes - node_filesystem_avail_bytes) / node_filesystem_size_bytes > 0.85
            for: 10m
            annotations:
              summary: 'Disk > 85%'

    Reference this file in prometheus.yml under rule_files: ['rules.yml'].

  6. Connect Grafana to Prometheus (and Loki)

    Open Grafana on port 3000. Default credentials: admin / admin — change them immediately.

    Add a Prometheus data source with URL http://prometheus:9090. Test the connection: it should return 'Data source is working'.

    If you added Loki to the Compose (grafana/loki:3.3.2), add a second Loki data source with URL http://loki:3100. You can then correlate CPU spikes in a Prometheus dashboard with application logs in Loki, without switching tools.

  7. Expose Grafana over HTTPS via a reverse proxy

    Grafana should never be exposed over plain HTTP. Configure Nginx as a reverse proxy:

    server {
      listen 443 ssl;
      server_name grafana.your-domain.com;
      ssl_certificate /etc/letsencrypt/live/grafana.your-domain.com/fullchain.pem;
      ssl_certificate_key /etc/letsencrypt/live/grafana.your-domain.com/privkey.pem;
      location / {
        proxy_pass http://127.0.0.1:3000;
        proxy_set_header Host $host;
      }
    }

    Block internal ports at the firewall: ufw deny 9090, ufw deny 9093, ufw deny 3000. Only port 443 (Nginx) should be publicly open.

  8. Verify that Prometheus is scraping all targets

    Go to http://localhost:9090/targets (via SSH or through the reverse proxy). All declared targets should show UP status in green. A DOWN status indicates a Docker network issue (wrong service name), a port issue, or a configuration problem. Fix it before moving on: a Grafana connected to a Prometheus with no data renders empty dashboards, and nothing in the interface explicitly says why.

Importing community Grafana dashboards

Grafana offers a library of ready-to-use dashboards at grafana.com/grafana/dashboards. Two essential identifiers for a VPS:

ID 1860 — Node Exporter Full: the reference dashboard for system metrics. It displays in real time and over time: CPU load per core, memory usage (including buffers/cache), disk I/O per device, network saturation and queue. Filterable by instance, making it usable to monitor multiple VPS from a single Grafana.

ID 893 — cAdvisor: per-container Docker metrics — CPU throttling, memory usage, I/O and network errors. Useful for detecting a container that consumes abnormally without slowing down the VPS overall.

To import a dashboard: in Grafana, go to Dashboards → Import, enter the ID, select your Prometheus data source and confirm. The dashboard is imported in seconds, with nothing to build by hand.

Configuring alerts with Alertmanager

Alertmanager is the component that receives alerts from Prometheus and decides what to do with them: group them, silence them during maintenance, route them to different channels based on severity.

Grouping: multiple alerts triggered at the same time (network outage → high CPU + inaccessible disk + service down) are grouped into a single notification, not three separate emails. This is the group_by: ['alertname', 'instance'] field in alertmanager.yml.

Silences: before planned maintenance, create a silence in the Alertmanager interface (http://localhost:9093) or via the API: amtool silence add --duration=2h alertname=~'.*'. Alerts continue to be evaluated by Prometheus but are not sent during the window.

Multi-channel routing: critical alerts (CPU > 95%) go to PagerDuty or Slack with @channel mention, warnings (CPU > 85%) to a less urgent monitoring channel. Configure this with nested routes and match_re in alertmanager.yml.

After each modification to alertmanager.yml, reload without restarting: curl -X POST http://localhost:9093/-/reload.

Prometheus vs InfluxDB vs Zabbix: which to choose?

Scroll the table

CriterionPrometheusInfluxDBZabbix
Collection modelPull (active scrape)Push (agents or API)Agent + SNMP + JMX
StorageBuilt-in TSDB, efficient on VPSInfluxDB OSS or CloudPostgreSQL / MySQL, heavy
Query languagePromQL (powerful, steep curve)Flux / InfluxQL (SQL-like)Zabbix macros (limited)
AlertingDedicated Alertmanager, flexibleBuilt-in (OSS limited)Built-in, rich but complex
Learning curveModerate (PromQL)Low (SQL-like)High (dense UI, agents)
Ideal use caseContainerized infra, instrumented appsIoT, high-frequency time seriesEnterprise networks, SNMP

Hardening: never expose Prometheus publicly

Prometheus has no native authentication system. Its web interface exposes the list of all your targets, your infrastructure labels and your alerting rules — valuable information for an attacker mapping your infrastructure.

Three non-negotiable rules:
1. Bind Prometheus on 127.0.0.1:9090 only (--web.listen-address=127.0.0.1:9090 in the Compose).
2. Block ports 9090, 9093, 9100 and 8080 at the firewall (ufw deny 9090).
3. Protect Grafana with a strong password and disable the anonymous account (GF_AUTH_ANONYMOUS_ENABLED=false as an environment variable).

If you need to access the Prometheus interface from outside for debugging, use an SSH tunnel (ssh -L 9090:localhost:9090 [email protected]) rather than a public reverse proxy.

Adding Docker container monitoring with cAdvisor

cAdvisor (Container Advisor) exposes per-container Docker metrics in Prometheus format: CPU usage as a proportion of the allocated quota, memory consumption with and without cache, network and disk I/O, and restarts.

In docker-compose.yml, add the service:

cadvisor:
  image: gcr.io/cadvisor/cadvisor:latest
  volumes:
    - /:/rootfs:ro
    - /var/run:/var/run:ro
    - /sys:/sys:ro
    - /var/lib/docker/:/var/lib/docker:ro
  ports:
    - '127.0.0.1:8080:8080'
  restart: unless-stopped

Then add the corresponding job in prometheus.yml. The Grafana dashboard ID 893 (cAdvisor) then displays a per-container view, with a filter on the Docker service name — useful for comparing consumption between your applications and detecting memory leaks.

Troubleshooting: the most common errors

Target DOWN in Prometheus (/targets): first check the service name in docker-compose.yml — Prometheus resolves names by Docker internal DNS. If the service is named node-exporter in the Compose, the scrape URL must be node-exporter:9100, not localhost:9100. Also verify that all services are on the same Docker network (networks: monitoring).

Empty dashboards after import: the most common cause is a scrape job whose label does not match what the dashboard expects. Node Exporter Full (ID 1860) expects a job named node — if you named it node_exporter in prometheus.yml, filter manually in the dashboard variables or rename the job.

Alertmanager not receiving alerts: verify that Prometheus can reach Alertmanager. In prometheus.yml, the alerting section must point to alertmanager:9093 (Docker service name). Test with curl http://alertmanager:9093/-/healthy from inside the Prometheus container: docker compose exec prometheus curl http://alertmanager:9093/-/healthy.

'context deadline exceeded' error on reload: Prometheus takes more than 30 s to reload if rules are numerous or if a target is slow. Increase the scrape timeout in prometheus.yml: scrape_timeout: 20s under global.

Going further

This stack covers the essentials of VPS monitoring. For more advanced needs:

Long-term retention: Prometheus stores data locally, ideal for 30 to 90 days. For a 1 to 2 year history without saturating the disk, connect Thanos (S3-compatible object store) or VictoriaMetrics as a drop-in storage replacement. VictoriaMetrics is particularly suited to modest VPS: it consumes 5 to 10× less memory than Prometheus for the same volume of series.

HTTP endpoint monitoring: add Blackbox Exporter to the Compose to test the availability of your URLs, the returned HTTP code, latency and TLS certificate expiry date — without instrumenting the application.

Logs with Loki: if you want to correlate metrics with application logs (Nginx errors, PHP exceptions...), add Loki + Promtail to the Compose. Grafana displays metrics and logs in the same panel, reducing diagnostic time.

For all advanced configuration options, refer to the official Grafana documentation and the Prometheus documentation. ServOrbit offers a pre-configured Docker template to launch this stack in one click on a Cloud VPS — Grafana, Prometheus, Node Exporter and Alertmanager mounted and ready, without initial configuration.

Deploy your observability stack

The ServOrbit Cloud VPS offers the CPU, RAM and SSD needed to run Prometheus, Grafana and their exporters without a bottleneck, with a Docker template ready to configure.

Need help?

Browse our help center and FAQ, or reach our team — callback, WhatsApp or email. Support in French, English and Arabic.

Message us on WhatsAppopens in a new tab