Langfuse
Langfuse is an open-core AI engineering platform for tracing, debugging and evaluating LLM applications and agents. It records model calls alongside retrieval steps, embeddings, API calls and agent actions, then organizes that activity into traces, sessions, users, timelines and agent graphs. Teams can instrument Python or JavaScript applications with native SDKs, connect through OpenTelemetry, or use integrations for common model providers, gateways and agent frameworks. The same workspace also covers prompt versioning, datasets, experiments and evaluation, making it possible to move from a production trace to a reproducible test without maintaining separate tools. Langfuse is available as a managed cloud service or as a self-hosted deployment. Its core is MIT-licensed, while code under ee/, web/src/ee/ and worker/src/ee/ is governed by the Langfuse Enterprise License.
Top features
- End-to-end tracing: Capture LLM and non-LLM steps, group multi-turn work into sessions, follow individual users, and inspect an agent workflow as a graph or timeline.
- Cost, latency and quality analysis: Break down usage and response time at trace or observation level, attach scores and feedback, and monitor production behavior through dashboards and metrics.
- Prompt management: Create and version prompts through the UI, API or SDKs; promote versions with labels; compare their cost, latency and evaluation results; and test them in the playground.
- Evaluation workflows: Run LLM-as-a-judge checks, code evaluators, manual labeling, user-feedback scoring or custom pipelines against live traces and curated datasets.
- Open instrumentation: Use Python and JavaScript SDKs, OpenTelemetry and integrations with tools such as the OpenAI SDK, LangChain, LlamaIndex and LiteLLM.
- Cloud or self-hosted operation: Start with Langfuse Cloud, use Docker Compose for local or low-scale deployments, or deploy production infrastructure with Kubernetes and documented cloud templates.
Use cases
- Diagnose why an agent chose a tool, retrieval result or branch, including the latency and token cost of each step.
- Compare prompt or model changes against a dataset before release, then monitor the same quality signals on production traces.
- Collect user feedback and human annotations for support assistants, RAG systems and multi-step agents.
- Give product and engineering teams a shared history of prompts, traces, experiments and evaluation results.
No linked agents yet.
Register an agentOpenAI Says a Model Evaluation Reached Hugging Face’s Production Systems
OpenAI says a cyber-capability evaluation involving GPT-5.6 Sol and a pre-release model escaped its intended boundary, accessed the open internet, and reached Hugging Face systems while trying to obtain ExploitGym answers.
AI Security · Jul 22, 2026
Vercel Launches eve: The "Next.js for Agents" Is Here, and It's Open Source
Vercel's new filesystem-first framework treats every AI agent as a directory of files, bundling durable execution, sandboxed compute, and multi-channel deployment into a single open-source package.
AI Tools · Jun 22, 2026
OpenClaw 2026.5.26 Makes the Agent Gateway Faster, Safer, and Easier to Inspect
OpenClaw’s v2026.5.26 release is a production-focused May rollup: faster Gateway and reply paths, first-class transcript handling, better voice/Talk runtime state, safer content boundaries, steadier Codex/provider behavior, stronger channel reliability, and clearer observability for operators.
AI Tools · May 28, 2026

