Why Replace Datadog or New Relic with SigNoz?
SaaS observability tools follow a punishing pricing model: free for a handful of hosts, then costs spiral with every new service. For a web agency managing ten client projects or an independent developer scaling a SaaS product, the monthly bill can easily exceed $200–500 with no proportional value gain.
SigNoz changes the equation. It is an open-source observability platform (Apache 2.0 license) powered by ClickHouse for high-performance trace and metric storage, and the OpenTelemetry protocol for instrumentation. Your data stays on your infrastructure, you control retention, and you only pay for the VPS running the stack.
What SigNoz Gives You
- Distributed traces: visualize the full path of an HTTP request across your microservices, with spans, durations, and errors at every hop.
- Prometheus-compatible metrics: import existing dashboards or build new ones directly in the SigNoz UI.
- Centralized logs: collect and correlate structured logs from all your applications in a unified interface.
- Configurable alerts: set thresholds on any metric and receive notifications via Slack, PagerDuty, or webhooks.
- Customizable dashboards: build business or technical views in a few clicks — no LoQL or PromQL required.
- Modern UI: responsive React interface accessible on port 8080 of your VPS, with no proprietary client-side agent.
Prerequisites Before You Start
SigNoz relies on ClickHouse, a columnar database engine that is memory-hungry. The recommended minimum is 4 GB of RAM — below that, ClickHouse gets killed by the kernel's OOM killer before the UI even loads. For production use with multiple instrumented applications, target 8 GB.
Full prerequisites:
- A VPS running Ubuntu 22.04 or Debian 12.
- Docker Engine ≥ 24 and Docker Compose V2 installed.
- Port 8080 open in your firewall (SigNoz UI).
- Ports 4317 (OTLP/gRPC) and 4318 (OTLP/HTTP) open to receive traces from your apps.
- Root or sudo access on the VPS.
- At least 20 GB of free disk space for ClickHouse and its data files.
Installing SigNoz via Foundry CLI
Step 1 — Install Docker on Your VPS
curl -fsSL https://get.docker.com | sh systemctl enable --now docker docker compose versionStep 2 — Install Foundry CLI (foundryctl)
Since v0.112.0, SigNoz uses the Foundry CLI as its official deployment method.
curl -L https://get.foundry.so/foundryctl/latest | bash export PATH="$HOME/.foundry/bin:$PATH" foundryctl --versionStep 3 — Create the casting.yaml File
mkdir -p /opt/signoz && cd /opt/signoz cat > casting.yaml << 'EOF' apiVersion: foundry.so/v1 kind: Casting metadata: name: signoz spec: release: stable components: - name: signoz enabled: true - name: clickhouse enabled: true EOFStep 4 — Run the Deployment
foundryctl cast -f casting.yamlFoundry CLI pulls the Docker images, sets up persistent volumes, and starts the containers in the correct order. Full startup takes about 2–3 minutes.
Step 5 — Verify SigNoz Is Running
docker compose -f /opt/signoz/docker-compose.yaml psThen open your browser at
http://<VPS_IP>:8080. Create your admin account on first login. Note: the old port 3301 mentioned in older tutorials is obsolete — the current UI port is 8080.Step 6 — Secure Access with a Reverse Proxy
Never expose port 8080 directly in production. Use Nginx as a reverse proxy with TLS:
server { listen 443 ssl; server_name signoz.yourdomain.com; ssl_certificate /etc/letsencrypt/live/signoz.yourdomain.com/fullchain.pem; ssl_certificate_key /etc/letsencrypt/live/signoz.yourdomain.com/privkey.pem; location / { proxy_pass http://127.0.0.1:8080; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; } }Block port 8080 from outside:
ufw deny 8080.
Instrumenting Node.js, Python and Go
SigNoz receives traces via the OTLP protocol. Here is how to wire up three common environments.
Node.js (Express/Fastify):
npm install @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-nodeOTEL_EXPORTER_OTLP_ENDPOINT="http://<VPS_IP>:4318" \
OTEL_SERVICE_NAME="my-api" \
node -r ./tracing.js app.jsPython (FastAPI/Flask):
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install
OTEL_EXPORTER_OTLP_ENDPOINT="http://<VPS_IP>:4318" \
OTEL_SERVICE_NAME="my-python-service" \
opentelemetry-instrument uvicorn main:appGo — manual spans:
go get go.opentelemetry.io/otel \
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp \
go.opentelemetry.io/otel/sdk/traceInitialize the tracer in your main.go, then create spans around critical operations:
tracer := otel.Tracer("my-go-service")
ctx, span := tracer.Start(ctx, "fetch-user-details")
defer span.End()
span.SetAttributes(attribute.String("user.id", userID))In all three cases, traces appear in the SigNoz Services view within seconds of the first HTTP call.
Configuring ClickHouse Retention and Cold Storage
On a VPS with limited disk space, the default SigNoz retention (3 days for traces, 30 days for metrics) can fill storage within weeks at moderate trace volume.
Adjust retention from the UI: go to Settings → Retention Period to set different durations for traces, metrics, and logs. For an agency managing multiple projects, 7 days of traces and 90 days of metrics strikes a good balance.
ClickHouse TTL with a cold volume:
If your VPS has a second, cheaper block volume mounted at /mnt/cold, you can configure ClickHouse to move aging data automatically:
<storage_configuration>
<disks>
<default/>
<cold_disk>
<type>local</type>
<path>/mnt/cold/clickhouse/</path>
</cold_disk>
</disks>
<policies>
<tiered>
<volumes>
<hot><disk>default</disk></hot>
<cold><disk>cold_disk</disk></cold>
</volumes>
</tiered>
</policies>
</storage_configuration>Then apply a TTL policy on the SigNoz traces table:
ALTER TABLE signoz_traces.distributed_signoz_index_v2
MODIFY TTL toDateTime(timestamp) + INTERVAL 7 DAY
TO VOLUME 'cold',
toDateTime(timestamp) + INTERVAL 30 DAY DELETE;ClickHouse moves data to the cold volume after 7 days and deletes it after 30 — no manual intervention required.
Custom Alerts in SigNoz
Once your applications are instrumented, SigNoz automatically populates the Services view with P50/P99 latency, error rate, and throughput for each service. You can alert on any of these signals.
Latency P99 alert: go to Alerts → New Alert Rule, choose Metric Based Alert, and enter the condition p99(signoz_latency_bucket{service_name="my-api"}) > 500.
Error rate alert: use rate(signoz_calls_total{service_name="my-api",status_code="STATUS_CODE_ERROR"}[5m]) > 0.05 to fire when more than 5% of calls fail over 5 minutes.
Anomaly detection: SigNoz supports Anomaly alert rules — define a reference window (e.g. 7 days) and an acceptable deviation threshold. Any unusual spike in throughput or latency triggers the alert — useful for catching slow degradations that fixed thresholds miss.
Custom dashboard: go to Dashboards → New Dashboard, add a Time Series panel, select a Prometheus metric, and apply filters by service or environment.
Scaling: Standalone vs Cluster
SigNoz in standalone mode (a single VPS, a single ClickHouse instance) handles up to roughly ten instrumented services generating a few thousand spans per second without effort. That is the target use case for a developer VPS.
When to consider a cluster:
- Your VPS regularly hits 80% CPU usage on the ClickHouse process.
- Trace ingestion exceeds 5,000 spans/s at sustained peak.
- You need high availability with zero downtime during updates.
Vertical scaling first: before distributing, double the VPS RAM. ClickHouse is designed to leverage large amounts of memory; moving from 8 to 16 GB RAM often shifts the bottleneck from CPU to network.
Horizontal scaling: SigNoz provides Helm charts for a Kubernetes deployment with distributed ClickHouse (shards + replicas). Reserve this for stacks that genuinely exceed standalone limits. On a VPS, the simplest approach is a second VPS dedicated to ClickHouse, connected to the first via a private network.
Watch Out for ClickHouse OOM Kills
ClickHouse is the most memory-intensive component in the SigNoz stack. On a VPS with less than 4 GB of available RAM, the Linux kernel may kill the ClickHouse process with exit code 137.
Diagnose it:
dmesg | grep -i oomYou will see a line like Out of memory: Killed process XXXX (clickhouse-serv).
Fixes: upgrade your VPS RAM (recommended), or cap ClickHouse memory by adding max_memory_usage=2000000000 to /etc/clickhouse-server/users.xml.
SigNoz vs Datadog vs Grafana Cloud
Scroll the table
| Criteria | SigNoz (self-hosted) | Datadog | Grafana Cloud |
|---|---|---|---|
| Monthly cost (5 services) | VPS cost only (~$10-20) | ~$75-150 + custom metrics | Free up to 10k series, then ~$8/1k |
| Distributed traces | Yes (native OpenTelemetry) | Yes (proprietary agent) | Yes (Tempo, via OTLP) |
| Metrics | Yes (Prometheus-compatible) | Yes (proprietary + Prometheus) | Yes (Mimir, Prometheus-compatible) |
| Logs | Yes (built-in) | Yes (additional cost) | Yes (Loki, additional cost) |
| Data sovereignty | Full — data on your VPS | Data at Datadog (US/EU) | Data at Grafana Labs |
| Operational complexity | Medium (Docker, 1 VPS) | None (SaaS) | Low (SaaS) |
| Updates | Manual (foundryctl) | Automatic | Automatic |
| Support | Community + paid plan | Paid (included) | Community + paid plan |
Deep Troubleshooting
UI does not load on port 8080: check that the signoz-frontend container is running and that your firewall allows the port.
Traces do not appear: the most common cause is an incorrect OTEL_EXPORTER_OTLP_ENDPOINT. Make sure the URL points to the public IP of the VPS (not localhost), uses port 4318 for OTLP/HTTP, and that no firewall blocks this port. Quick test from your application server:
curl -v http://<VPS_IP>:4318A 405 Method Not Allowed response confirms the collector is listening.
ClickHouse OOM — container restart loop: confirm with dmesg | grep -i oom. If RAM is insufficient, cap ClickHouse memory usage:
echo '<max_memory_usage>2000000000</max_memory_usage>' \
>> /etc/clickhouse-server/users.xml
docker compose restart clickhouseSigNoz startup timeout: if docker compose ps shows the signoz container stuck in starting for more than 5 minutes, ClickHouse is not yet ready. SigNoz waits for a ClickHouse response before initializing. Check ClickHouse container logs:
docker compose logs clickhouse --tail=50A full disk or insufficient permissions on the data volume are the most common causes.