Why LLM applications need dedicated observability
Traditional APM tools (Datadog, New Relic) capture HTTP latency and error rates — but they are blind to what happens inside an LLM call. A response that arrives in 800 ms might still be factually wrong, unhelpfully vague, or three times more expensive than yesterday's because a prompt regression slipped through code review. Langfuse solves this by treating each LLM interaction as a structured trace: it records the full prompt (including system message and conversation history), the completion, the model used, the token counts, the latency breakdown per span, and any evaluation scores your team attaches. You can then filter, compare, and reproduce any trace — individually or in aggregate.
What Langfuse gives you out of the box
- Full-stack LLM tracing: prompt, completion, latency, cost and token count — including nested spans for agent chains (LangChain, LlamaIndex, Dify).
- Prompt management hub: version-control your prompts, stage variants, A/B test in production, promote the winner without a code deploy.
- Evaluation framework: run LLM-as-a-judge, human annotation queues, or custom scoring functions on any trace or dataset.
- Cost analytics: track token spend by model, endpoint, user and session — switch providers with data, not guesswork.
- Dataset management: capture production traces as golden test sets for offline evaluation and regression detection.
- Native SDK for Python and TypeScript, plus automatic integration with LiteLLM, LangChain, LlamaIndex, Dify, Haystack and VercelAI.
What Langfuse runs on (and why it needs 4 GB RAM)
Langfuse v3 ships as a six-service Docker Compose stack: langfuse (Next.js frontend + API), langfuse-worker (background jobs and evaluations), postgres (application state), clickhouse (trace analytics — columnar storage optimised for high-cardinality time-series), redis (queue and cache), and minio (S3-compatible blob storage for media attachments). ClickHouse is the reason for the 4 GB minimum: its JIT compiler and vector execution engine need headroom. On a quiet VPS the full stack idles around 1.5–2 GB; under moderate production load plan for 4 GB, with 8 GB for sustained high-throughput tracing. All six services start from a single docker compose up -d and are managed by AWX in the ServOrbit one-click flow.
Deploy Langfuse on your VPS in six steps
Order a VPS Power (8 GB RAM)
Langfuse needs at least 4 GB of RAM: in the catalogue the first plan that clears that bar is the VPS Power (4 vCPU, 8 GB), installed on Ubuntu 24.04 — enough to absorb persistent production traffic too. Pick a domain or subdomain you control — Langfuse's NextAuth session cookies require a proper HTTPS domain.
One-click install from the marketplace
Open your ServOrbit control panel, go to Marketplace → Artificial Intelligence → Langfuse, and click Deploy. Enter your domain when prompted. Docker Compose pulls all six images and starts them; the web UI is ready on port 3000 within 60–90 seconds (ClickHouse first-boot initialisation takes the longest).
Open the web UI and create your first project
Navigate to
https://your-domain.com. Langfuse shows the signup screen on first boot. Create your admin account, then go to Settings → Projects → Create project. Copy the Public Key and Secret Key from the project settings — you'll need them in your application.Instrument your Python or TypeScript application
In Python:
pip install langfuse, setLANGFUSE_HOST=https://your-domain.com,LANGFUSE_PUBLIC_KEYandLANGFUSE_SECRET_KEY, then decorate your LLM calls with@observe()or uselangfuse.trace(). In TypeScript:npm install langfuse, initialise the client with your host and keys. Traces appear in the dashboard within seconds of the first call.Enable automatic tracing if you use LiteLLM
If LiteLLM is already in your stack (which it likely is if you use the ServOrbit AI stack), add
success_callback = ["langfuse"]tolitellm_config.yamland set the three Langfuse env vars. Every proxied LLM call — across all downstream models — is traced automatically, with cost and token data attached.Run your first evaluation
Go to Traces, filter for a representative sample, click 'Add to dataset'. Open Datasets → your dataset → Run evaluation, choose LLM-as-a-judge with a prompt template ('Rate the relevance of this response on a scale of 1–5'), and submit. Scores appear on every trace in the set and feed into the aggregate analytics dashboard.
Logging in for the first time
The URL opens the Langfuse screen with a "Sign up" link: create your account, then your organisation and your first project in the wizard.
Post-install configuration: nginx reverse proxy, TLS and key environment variables
The ServOrbit Marketplace installation configures nginx and the TLS certificate for you. If you are setting up manually, here are the critical points.
nginx reverse proxy. Langfuse listens on port 3000. Your nginx vhost must proxy this port and forward the headers NextAuth requires:
server {
listen 443 ssl;
server_name your-domain.com;
location / {
proxy_pass http://127.0.0.1:3000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}If X-Forwarded-Proto is not forwarded, NextAuth generates callback URLs over http:// and session cookies do not attach on the HTTPS response — symptom: redirect loop after login.
Key environment variables. The Langfuse stack .env file has three variables you must set before the first start:
- NEXTAUTH_SECRET: a long random string (minimum 32 characters). Without it, sessions do not persist across restarts.
- NEXTAUTH_URL: the full public URL of your instance (https://your-domain.com). Must exactly match the nginx domain.
- LANGFUSE_INIT_ORG_ID, LANGFUSE_INIT_PROJECT_ID, LANGFUSE_INIT_PROJECT_PUBLIC_KEY, LANGFUSE_INIT_PROJECT_SECRET_KEY: optional, but useful for provisioning an organisation and project on first boot without using the UI.
TLS. ServOrbit deployments use Let's Encrypt via certbot in DNS-01 mode — no port 80 required. If you configure the certificate manually, point the ssl_certificate and ssl_certificate_key directives at the certbot-issued files and enable automatic renewal (certbot renew --quiet in cron).
Pair with LiteLLM for complete AI cost visibility
The ServOrbit AI stack already includes LiteLLM as an OpenAI-compatible gateway. Connect Langfuse to LiteLLM with three env vars and you get a complete picture: LiteLLM enforces rate limits and routes between providers; Langfuse records every trace with full prompt, completion, model, tokens and cost. Together they give you a private, auditable AI operations layer — no SaaS middleman, no data egress.
SDK integration: tracing your LLM agents in Python and TypeScript
Langfuse ships two official SDKs covering the main agent frameworks.
Python — @observe() decorator.
The most direct way to trace an agent:
from langfuse.decorators import observe, langfuse_context
@observe()
def run_agent(user_message: str) -> str:
# every nested LLM call is automatically captured
response = openai_client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": user_message}]
)
return response.choices[0].message.contentAll nested functions decorated with @observe() appear as child spans within the same trace. To attach metadata (user ID, session tags):
langfuse_context.update_current_trace(
user_id="user-42",
session_id="session-abc",
tags=["production", "agent-v2"]
)LangChain. The native integration activates in one line:
from langfuse.callback import CallbackHandler
handler = CallbackHandler() # reads LANGFUSE_* from env
chain.invoke({"input": query}, config={"callbacks": [handler]})Every node in a LangGraph graph (or LangChain chain) becomes a span with its own token count and cost.
TypeScript / JavaScript.
Client initialisation:
import Langfuse from "langfuse";
const lf = new Langfuse({
publicKey: process.env.LANGFUSE_PUBLIC_KEY!,
secretKey: process.env.LANGFUSE_SECRET_KEY!,
baseUrl: process.env.LANGFUSE_HOST,
});Creating a trace and a generation:
const trace = lf.trace({ name: "agent-run", userId: "user-42" });
const generation = trace.generation({
name: "llm-call",
model: "gpt-4o",
input: messages,
});
// … LLM call …
generation.end({ output: result, usage: { input: 120, output: 45 } });
await lf.shutdownAsync(); // flush buffer before process exitsVercel AI SDK. If you use the Vercel AI SDK, the Langfuse integration activates by adding langfuseExporter as a SpanExporter in the OpenTelemetry configuration — all generations are traced without changing your business logic.
LLM costs showing $0.00 in Tracing v4 — what is happening and how to work around it
An active bug affects Langfuse's Tracing v4 view (tracked in issue #16077, 17 comments in August 2026): the Cost ($) column always displays $0.00 for every trace.
The root cause is structural. The v4 view reads the cost stored on the root observation of each trace, which is of type SPAN — a span carries no cost of its own; it only wraps child generation observations. The actual cost is distributed across child GENERATION observations, but the view does not aggregate them: it reads the root, finds zero, and displays zero.
Your data is intact. Cost is correctly recorded on each generation; only the display in the main view is wrong. Three workarounds are available while the official fix is pending:
1. Analytics / Metrics tab — the aggregate view computes costs across all generations, not just the root. This is the reliable source for tracking your daily spend.
2. Filter by GENERATION type in Traces — add the filter observation_type = GENERATION. Every row shows its real cost.
3. Open the trace detail panel — the side panel shows the full breakdown by span and generation with correct costs. Useful to diagnose a specific trace.
If you observe this on your instance, confirm you are running Langfuse v3.x with Tracing v4 enabled. The issue is open and tracked by the Langfuse team.
Troubleshooting: five common errors after installation
1. The langfuse container does not start — Error: NEXTAUTH_SECRET missing.
Langfuse v3 refuses to start without this variable. Check your .env:
docker compose config | grep NEXTAUTH_SECRETIf the line is absent or empty, generate a value and restart:
openssl rand -base64 32
# paste the result into .env: NEXTAUTH_SECRET=<value>
docker compose up -d langfuse2. Database connection refused — ECONNREFUSED postgres:5432.
This is almost always a startup timing issue: langfuse attempts to connect before postgres is ready. Make sure your docker-compose.yml declares depends_on with condition: service_healthy for the postgres service. If the problem persists, inspect the postgres logs:
docker compose logs postgres --tail=503. Traces are not appearing in the dashboard.
Check in order: (a) are the environment variables LANGFUSE_HOST, LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY set in your application? (b) Does the host point to the public HTTPS URL and not to localhost:3000? (c) Is the Python SDK flushing its buffer before the process exits (langfuse.flush() or langfuse.shutdown())?
A quick test call from Python:
from langfuse import Langfuse
lf = Langfuse()
lf.trace(name="test-connection")
lf.flush()
print("Trace sent — check the dashboard.")4. Redirect loop after login.
Most common cause: NEXTAUTH_URL does not match the domain in use, or X-Forwarded-Proto is not forwarded by nginx. Verify that NEXTAUTH_URL=https://your-domain.com (no trailing slash) and that your nginx config includes proxy_set_header X-Forwarded-Proto $scheme;.
5. ClickHouse consuming all available RAM.
Under heavy load, ClickHouse can exceed its initial allocation. Limit memory consumption by adding to your .env:
CLICKHOUSE_MAX_MEMORY_USAGE=2000000000This value (2 GB) is a reasonable starting point for an 8 GB VPS running only Langfuse. Adjust based on your trace volume.
Langfuse vs Langsmith vs Helicone
Langsmith (LangChain's hosted product) and Helicone are the main SaaS alternatives. Both are excellent — and both require you to route your traces through their servers. For teams handling sensitive prompts (legal, financial, medical), or for compliance environments that prohibit third-party data egress, self-hosting is not optional. Langfuse gives you the same feature set (tracing, evals, prompt management, cost analytics) on infrastructure you control. The MIT licence also means you can read, modify and audit every line of the code your traces flow through.