SKILL PROCEDURE

Langfuse

Use when instrumenting, debugging, or evaluating an LLM application with Langfuse — tracing calls and agent steps, running LLM-as-a-judge or human evaluations, managing and versioning prompts, or deciding between Langfuse Cloud and a self-hosted deployment. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.

langfuseobservabilitytracingevaluationprompt-management
BEGINNER GUIDE

Understand Langfuse before using it

CATEGORY

Langfuse is catalogued under AI and ML infrastructure.

START HERE WHEN

Your work repeatedly involves the concepts tagged above. Open the full procedure below when the current task matches them.

Compare related skills

The current Hardgraph catalogue has no close alternative with the same category or concepts. Read the full procedure rather than comparing it with an unrelated tool.

Langfuse

What is HardGraph? HardGraph publishes curated, provenance-backed agent skills grounded in reproducible vendor documentation.

Langfuse is an open-source LLM engineering platform built around one data model — the trace — that three largely independent products consume: tracing/observability, evaluation, and prompt management. Treat them as separate concerns even when reaching for all three at once: a trace is captured regardless of whether anything evaluates or prompts against it, an evaluation can run against traces captured weeks earlier by a different SDK, and prompt management works even for an application that sends Langfuse no traces at all.

The decision that shapes everything else

Cloud versus self-hosted is a data-residency decision, not a convenience one. Langfuse Cloud means every trace — prompts, completions, and any metadata or user identifiers you attach — leaves your infrastructure and lands on Langfuse's servers, subject to their region and retention settings. Self-hosting keeps that data in infrastructure you control, but it is not a single container: a production deployment coordinates a relational store, an analytical store for trace data, a cache, and blob storage, each independently upgraded and scaled. Decide this before instrumenting anything — traces don't migrate between deployment modes for free, and masking/PII-scrubbing decisions are cheaper to make before the first trace ships than after.

Easily confused, worth getting right

  • Tracing is not evaluation. Tracing answers "what happened"; evaluation (LLM-as-a-judge, human annotation, code-based scorers) answers "was it good," and runs as a separate pass — sometimes online against live traffic, sometimes offline against a fixed dataset. Conflating them leads to instrumenting for observability and then being surprised evaluation needs its own scorer configuration.
  • Prompt management is not prompt engineering inside your codebase. Prompts pulled from Langfuse are versioned and can be updated without a deploy — which also means an application can silently start running a different prompt than the one committed in its repo.
  • The ingestion protocol has moved across major versions, including a shift toward OpenTelemetry as the transport. An SDK and a server instance from mismatched eras can fail to talk to each other in ways that look like a networking bug.

What to verify rather than recall

Do not assert a remembered SDK version, self-hosted infrastructure requirement, or which features require an enterprise license key — these have changed across major versions and will again. Verify current guidance, SDK migration paths, and self-hosted component requirements against Langfuse's own documentation before committing to an approach.

Hardgraph / curated knowledge for agents.

STATIC EXPORT · CANONICAL SOURCE