Langfuse
Open-source LLM engineering platform unifying tracing, evaluation, prompt management, and metrics dashboards; MIT-licensed, self-hostable, and used by 2,300+ companies processing billions of observations a month.
Visit siteA curated 2026 directory of 14 platforms for tracing, evaluating, and monitoring large language models and AI agents in production — from open-source tracing libraries to enterprise-grade, GPU-to-guardrail observability suites.
Prompt-to-response and tool-call chains
LLM-as-judge, quality & regression scoring
Cost, latency, and token usage
Safety, PII, and compliance checks
End-to-end platforms combining tracing, evaluation, prompt management, and production monitoring in one product.
Open-source LLM engineering platform unifying tracing, evaluation, prompt management, and metrics dashboards; MIT-licensed, self-hostable, and used by 2,300+ companies processing billions of observations a month.
Visit siteEnd-to-end observability and evals for the agentic era — traces every agent step, scores outputs with 30+ metrics, and includes "Ollie," a coding assistant that writes fixes directly into your codebase.
Visit siteActive observability platform for shipping quality agents — real-time trace inspection, LLM/code/human scoring, and automatic pattern discovery, backed by Brainstore, a database purpose-built for agent traces.
Visit siteLangChain's agent and LLM observability platform — framework-agnostic tracing, real-time monitoring dashboards, and automatic insight clustering across production traces, powered by the purpose-built SmithDB.
Visit siteLLM and agent observability extended from established enterprise APM and infrastructure-monitoring suites — full-stack visibility from GPU to prompt.
Full-stack monitoring for GenAI apps, LLMs, and agentic workflows — covering orchestration, model integrity, guardrails, and GPU/infrastructure layers within Dynatrace's broader observability suite.
Visit siteExtends Datadog's APM stack to LLM calls and agent chains — end-to-end tracing, quality monitoring, and cost visibility alongside the infrastructure and app metrics teams already track in Datadog.
Visit siteElasticsearch-powered LLM monitoring with prebuilt dashboards for Bedrock, Azure AI Foundry, OpenAI, and Anthropic — plus guardrail tracking and OpenTelemetry-based step-by-step tracing.
Visit siteCisco-Splunk platform (built on the Galileo acquisition) that evaluates, observes, and guards AI agents end-to-end, correlating quality signals with GPU-level infrastructure data and per-agent token cost ("Tokenomics").
Visit siteObservability platform built for AI-era software — agent timeline views and LLM observability sit on top of a fast, OpenTelemetry-native query engine with BubbleUp root-cause analysis.
Visit siteDeveloper-first, open-source tools focused on instrumentation, trace visualization, and rigorous, code-native evaluation.
Open-source OpenTelemetry instrumentation for LLM applications — add two lines of code and stream traces to Dynatrace, Datadog, Honeycomb, Splunk, New Relic, or any OTel-compatible backend.
Visit siteOpen-source AI development platform for tracing, annotating, and evaluating agents, with an in-app AI engineering agent ("PXI") and one-click self-hosting via terminal, Docker, or Kubernetes.
Visit siteOpen-source, pytest-native LLM evaluation framework with 50+ research-backed metrics for hallucination, faithfulness, and agent task completion — runnable in CI/CD or straight from a coding agent's terminal.
Visit siteGateway-layer tools that route and cache LLM calls across providers while logging every request for cost, latency, and quality analysis.
AI gateway and observability layer for routing, caching, and monitoring LLM calls across 250+ providers, with unified request logs, guardrails, and cost/latency dashboards from a single API.
Visit siteOpen-source AI gateway and LLM observability platform for routing, debugging, and analyzing production AI applications, with request-level tracing, sessions, prompts, and dataset tooling.
Visit site