Why choose OpenTelemetry over a proprietary agent
OpenTelemetry is a graduated CNCF project that unifies the three pillars of observability — traces, metrics, and logs — under a vendor-neutral standard. Unlike Datadog or New Relic agents that lock you into a proprietary contract and format, the OTel Collector receives, transforms, and exports telemetry to any compatible backend. You can switch storage without changing a single line of application code.
SDKs cover Node.js, Python, Go, Java, PHP, Ruby, and most modern languages. W3C TraceContext and B3 context propagators are included: they automatically forward the trace-id between your microservices via HTTP headers, with no extra configuration. The OTel Collector can run as a local agent on each VM or as a centralized gateway — you choose the topology that fits your infrastructure.
What the OTel + Grafana Tempo stack gives you
- Distributed tracing without an external index — Grafana Tempo v2.x stores traces in compressed blocks (parquet) on local disk or object storage, with no Cassandra or Elasticsearch.
- TraceQL queries — Tempo's dedicated language lets you filter traces by duration, service, error status, or custom attributes in seconds.
- Multi-protocol compatibility — Tempo accepts OTLP, Zipkin, and Jaeger: your existing applications connect without full re-instrumentation.
- Lightweight on VPS — 2 GB of RAM is sufficient for a dev or staging environment; in production under load, 4 GB provides a comfortable margin.
- Controlled cost — a self-hosted VPS at 99 DH/month replaces a Datadog APM subscription billed at around $23 per host per month.
- Traces–logs–metrics correlation — by linking Tempo to Loki and Prometheus in Grafana, you jump from a trace to its logs and metrics in a single click.
- No vendor lock-in — if you migrate to Jaeger tomorrow, you reconfigure one exporter in the Collector; your SDKs and application code remain untouched.
VPS prerequisites
Your VPS needs at least 2 GB of RAM (4 GB recommended for production) and root or sudo access. Docker Engine ≥ 24 and Docker Compose v2 must be installed.
Plan a dedicated subdomain — for example grafana.your-domain.com — to expose the Grafana interface behind a TLS reverse proxy (nginx or Caddy). Ports 4317 (OTLP/gRPC) and 4318 (OTLP/HTTP) must be open in your firewall on the private side to receive telemetry from your applications; do not expose them on the public interface. Tempo port 3200 and Grafana port 3000 remain internal to the Docker network.
For storage, plan for at least 10 GB of free space for trace blocks. The default mount at /tmp/tempo does not survive restarts: replace it with a named Docker volume or a dedicated disk mount before going to production.
Stack architecture: from span to visualization
The data flow follows four steps:
1. Your application instruments its requests with an OTel SDK and exports spans to the OTel Collector (OTLP gRPC port 4317 or HTTP port 4318).
2. The OTel Collector receives, batches, then exports to Tempo via internal OTLP gRPC. It can simultaneously send metrics to Prometheus and logs to Loki.
3. Grafana Tempo stores compressed trace blocks on disk (or S3). It exposes a TraceQL API on port 3200 for Grafana queries.
4. Grafana queries Tempo with TraceQL, displays flamegraphs, service graphs, and span metrics. If Tempo is linked to Loki, clicking a span opens the logs for that service on the same time window.
TraceQL is Tempo's native query language. A typical query looks like: {service.name="api" && status=error && duration>500ms}. You filter by resource attribute, span attribute, status, or duration — and immediately get traces that cross a latency threshold.
Span Metrics is a Tempo pipeline that automatically derives RED metrics (Rate, Error, Duration) from incoming traces, without adding a single counter to your code. These metrics can be scraped by Prometheus and displayed in standard Grafana dashboards.
Deploy the OTel + Tempo stack in seven steps
Create the directory structure
Set up an
otel-stack/folder on your VPS:mkdir -p otel-stack/config cd otel-stackYou will place three configuration files here:
tempo.yaml,otel-collector-config.yaml, anddocker-compose.yml.Write the Tempo configuration (tempo.yaml)
Create
config/tempo.yamlwith local storage backend and TraceQL support:server: http_listen_port: 3200 distributor: receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 ingester: max_block_bytes: 1_000_000 max_block_duration: 5m compactor: compaction: block_retention: 168h # 7 days storage: trace: backend: local wal: path: /tmp/tempo/wal local: path: /tmp/tempo/blocksAdjust
block_retentionto match your desired retention period.Write the OTel Collector configuration
Create
config/otel-collector-config.yaml:receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318 processors: batch: timeout: 5s send_batch_size: 1000 exporters: otlp: endpoint: tempo:4317 tls: insecure: true service: pipelines: traces: receivers: [otlp] processors: [batch] exporters: [otlp]Write the docker-compose.yml
Create
docker-compose.ymlat the root ofotel-stack/:services: otel-collector: image: otel/opentelemetry-collector-contrib:latest volumes: - ./config/otel-collector-config.yaml:/etc/otelcol-contrib/config.yaml ports: - "4317:4317" - "4318:4318" networks: - observability depends_on: - tempo tempo: image: grafana/tempo:latest command: ["-config.file=/etc/tempo.yaml"] volumes: - ./config/tempo.yaml:/etc/tempo.yaml - tempo-data:/tmp/tempo ports: - "3200:3200" networks: - observability grafana: image: grafana/grafana:latest ports: - "3000:3000" environment: - GF_AUTH_ANONYMOUS_ENABLED=true - GF_AUTH_ANONYMOUS_ORG_ROLE=Admin networks: - observability depends_on: - tempo volumes: tempo-data: networks: observability:Start the stack:
docker compose up -dthen verify all three services are healthy withdocker compose ps.Configure Grafana Tempo as a datasource
Open Grafana at
http://your-vps:3000. Go to Connections → Data sources → Add data source and select Tempo. Set the URL tohttp://tempo:3200. Enable Service graph and Span metrics to enrich the visualization with dependency graphs and automatically computed RED metrics. If your stack includes Loki, link both datasources in the Linked data sources field: you will be able to navigate from a span to its logs in one click.Send your first test trace
Test that the stack receives traces using
telemetrygen:docker run --rm --network otel-stack_observability \ ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest \ traces \ --otlp-endpoint otel-collector:4317 \ --otlp-insecure \ --duration 5s \ --rate 10In Grafana, open Explore, select the Tempo datasource, and run the query
{}to see all received traces. If no traces appear, check the Collector logs:docker compose logs otel-collector.Query your traces with TraceQL
In the Grafana Explore panel (Tempo datasource), enter a TraceQL query to filter your traces. Useful examples:
-
{service.name="my-api"}— all traces from themy-apiservice
-{status=error}— all traces in error state
-{duration>1s && service.name="my-api"}— slow traces taking more than one second
-{span.http.route="/api/orders" && status=error}— errors on a specific routeClick on a trace to open the detailed flamegraph: each span shows its duration, attributes, and any errors.
Configure trace retention and storage
By default, Tempo stores blocks locally in /tmp/tempo. For a durable installation, two options are available.
Local storage on a Docker volume: this is the recommended configuration for a standalone VPS. The tempo-data volume defined in docker-compose.yml persists data across restarts. Adjust block_retention in tempo.yaml according to your needs: 72 h for a staging environment, 168 h (7 days) to 720 h (30 days) for production. Each hour of compressed traces occupies approximately 100–300 MB depending on span volume.
S3-compatible object storage: for a multi-node stack or uncapped storage, replace the local backend with an s3 backend in tempo.yaml:
storage:
trace:
backend: s3
s3:
bucket: my-traces-bucket
endpoint: s3.amazonaws.com
region: eu-west-1
access_key: ${S3_ACCESS_KEY}
secret_key: ${S3_SECRET_KEY}
wal:
path: /tmp/tempo/walTempo is compatible with AWS S3, MinIO, Scaleway Object Storage, and any S3 API-compatible backend. The WAL (Write-Ahead Log) always stays on local disk to ensure durability of in-flight traces.
Tempo's compactor merges small blocks and enforces the retention policy. It runs in the background and requires no manual intervention. Monitor the blocks/ folder size with docker exec tempo du -sh /tmp/tempo/blocks to anticipate storage needs.
Instrument a Node.js application
OTel auto-instrumentation for Node.js automatically captures outgoing HTTP calls, Express/Fastify requests, database queries, and many other libraries without modifying your business logic.
Install the required packages:
npm install @opentelemetry/sdk-node \
@opentelemetry/auto-instrumentations-node \
@opentelemetry/exporter-trace-otlp-grpcCreate a tracing.js file to load first:
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || 'grpc://localhost:4317',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();Start your application with this file loaded first:
export OTEL_SERVICE_NAME=my-api
export OTEL_EXPORTER_OTLP_ENDPOINT=http://your-vps:4317
node --require ./tracing.js app.jsFor Docker environments, pass the OTEL_SERVICE_NAME and OTEL_EXPORTER_OTLP_ENDPOINT variables in your docker-compose.yml or Kubernetes manifest. All incoming and outgoing HTTP calls will be automatically traced and correlated by the W3C TraceContext propagator.
Instrument a Python application
The opentelemetry-distro package provides a bootstrap mechanism that automatically installs all instrumentation packages for libraries detected in your environment.
Install and bootstrap:
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a installLaunch your application prefixed with opentelemetry-instrument:
export OTEL_SERVICE_NAME=my-python-service
export OTEL_EXPORTER_OTLP_ENDPOINT=http://your-vps:4317
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
opentelemetry-instrument python app.pyDjango, Flask, FastAPI, SQLAlchemy, and HTTP clients (requests, httpx, aiohttp) are instrumented automatically. For containerized deployments, define the OTEL_* variables in your environment file and ensure port 4317 of the OTel Collector is accessible from your containers' network.
For manual instrumentation with finer control:
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
with tracer.start_as_current_span("process-order") as span:
span.set_attribute("order.id", order_id)
# your business logicOTel + Tempo vs Jaeger vs Zipkin vs Datadog APM
Scroll the table
| Criterion | OTel + Grafana Tempo | Jaeger | Datadog APM |
|---|---|---|---|
| Monthly cost | Hosting only (99 DH/month VPS) | Hosting only (similar) | ~$23 per host per month |
| Storage backend | Local, S3, GCS, Azure Blob | Cassandra, Elasticsearch, Badger | Managed SaaS (opaque) |
| Query language | TraceQL (native, powerful) | Jaeger Query UI (basic) | APM Query Language (proprietary) |
| Auto-instrumentation | Universal OTel SDK | Jaeger libraries or OTel | Proprietary Datadog agent |
| Traces–logs correlation | Yes (Loki + Grafana) | Limited | Yes (integrated, paid) |
| Span Metrics (RED) | Yes (native Tempo pipeline) | Not native | Yes (computed by Datadog) |
| Vendor lock-in | None (CNCF standard) | Low (OTel compatible) | Strong (proprietary agent and format) |
| Installation complexity | Medium (3 containers) | Low to medium | Low (auto agent) |
Alerts and SLOs in Grafana
Grafana 10+ lets you define alerts directly on metrics derived from traces. Tempo's Span Metrics automatically generate three metrics per service: traces_spanmetrics_calls_total (request rate), traces_spanmetrics_latency_bucket (latency histogram), and the associated error counter.
To create an SLO on your API error rate:
1. In Grafana, go to Alerting → Alert rules → New alert rule.
2. Choose the Prometheus datasource (or Tempo via Span Metrics) and enter a query like rate(traces_spanmetrics_calls_total{status_code="STATUS_CODE_ERROR", service="my-api"}[5m]) / rate(traces_spanmetrics_calls_total{service="my-api"}[5m]).
3. Set the alert threshold (for example > 0.05 for an error rate above 5%) and the evaluation window.
4. Configure a contact point: Alertmanager, PagerDuty, Slack, or webhook — all integrate from Alerting → Contact points.
For P99 latency, query histogram_quantile(0.99, rate(traces_spanmetrics_latency_bucket{service="my-api"}[5m])) and trigger an alert if the value exceeds your error budget (for example 2 seconds). These trace-based alerts are far more precise than synthetic metric alerts, as they reflect the real experience of each individual request.
Combine traces, logs, and metrics in a single Collector
The OTel Collector is not limited to traces. By adding a filelog receiver to your configuration, you collect logs from your Docker containers and ship them to Loki. A prometheus receiver scrapes your /metrics endpoints and exports them to Mimir or Prometheus. A single process handles all telemetry from your VPS.
In Grafana, link Tempo and Loki via derived fields: when a log contains a trace_id, a clickable link takes you directly to the corresponding flamegraph. Conversely, from a Tempo trace, a link opens the Loki logs for that service in the associated time window. This cross-navigation dramatically reduces diagnosis time: no more manually correlating timestamps across three separate tabs.
Troubleshooting common issues
No spans arriving in Tempo. First check that the OTel Collector is reachable from your application (telnet your-vps 4317). Check the Collector logs (docker compose logs otel-collector): a connection refused error on tempo:4317 indicates Tempo is not yet started or its internal port is not exposed on the Docker network. Ensure both services share the same network (observability).
Connection refused from application to Collector. On a UFW or firewalld firewall, allow ports 4317 and 4318 only for your application IPs, not for the entire internet. In production, applications and Collector should be on the same private network; never expose the Collector on the public interface without authentication.
Truncated traces or missing spans. The Collector's batch processor groups spans before sending. If your application sends few spans, traces can take up to timeout seconds (5 s by default) to appear in Grafana. For development, reduce timeout: 1s in the batch processor. If spans are systematically absent, verify the context propagator is activated in your SDK (W3CTraceContextPropagator) — without it, calls between services generate orphaned traces.
Tempo consuming too much disk space. Reduce block_retention in tempo.yaml and restart Tempo. The compactor will apply the new policy on the next cycle (every 5 minutes by default). In an emergency, manually delete old blocks in /tmp/tempo/blocks while leaving the WAL intact.
Grafana can't find traces from Loki logs. Verify that your applications inject trace_id into their logs (structured JSON format recommended) and that the Derived fields setting in the Loki datasource points to the correct pattern ("traceId": "(\w+)"). The link URL must point to the Tempo datasource with the correct Grafana datasource identifier.