Tracing · Evaluation · Monitoring · AI Gateways

LLM Observability Platforms & Tools

A curated 2026 directory of 14 platforms for tracing, evaluating, and monitoring large language models and AI agents in production — from open-source tracing libraries to enterprise-grade, GPU-to-guardrail observability suites.

🔍

Tracing

Prompt-to-response and tool-call chains

🧪

Evaluation

LLM-as-judge, quality & regression scoring

📊

Metrics

Cost, latency, and token usage

🛡️

Guardrails

Safety, PII, and compliance checks

Full Observability Platforms

4 tools

End-to-end platforms combining tracing, evaluation, prompt management, and production monitoring in one product.

01Full Platform

Langfuse

Open-source LLM engineering platform unifying tracing, evaluation, prompt management, and metrics dashboards; MIT-licensed, self-hostable, and used by 2,300+ companies processing billions of observations a month.

Visit site
02Full Platform

Opik by Comet

End-to-end observability and evals for the agentic era — traces every agent step, scores outputs with 30+ metrics, and includes "Ollie," a coding assistant that writes fixes directly into your codebase.

Visit site
03Full Platform

Braintrust

Active observability platform for shipping quality agents — real-time trace inspection, LLM/code/human scoring, and automatic pattern discovery, backed by Brainstore, a database purpose-built for agent traces.

Visit site
04Full Platform

LangSmith Observability

LangChain's agent and LLM observability platform — framework-agnostic tracing, real-time monitoring dashboards, and automatic insight clustering across production traces, powered by the purpose-built SmithDB.

Visit site

Enterprise & APM-Native Observability

5 tools

LLM and agent observability extended from established enterprise APM and infrastructure-monitoring suites — full-stack visibility from GPU to prompt.

05Enterprise

Dynatrace AI Observability

Full-stack monitoring for GenAI apps, LLMs, and agentic workflows — covering orchestration, model integrity, guardrails, and GPU/infrastructure layers within Dynatrace's broader observability suite.

Visit site
06Enterprise

Datadog LLM Observability

Extends Datadog's APM stack to LLM calls and agent chains — end-to-end tracing, quality monitoring, and cost visibility alongside the infrastructure and app metrics teams already track in Datadog.

Visit site
07Enterprise

Elastic LLM Observability

Elasticsearch-powered LLM monitoring with prebuilt dashboards for Bedrock, Azure AI Foundry, OpenAI, and Anthropic — plus guardrail tracking and OpenTelemetry-based step-by-step tracing.

Visit site
08Enterprise

Splunk Agent Observability

Cisco-Splunk platform (built on the Galileo acquisition) that evaluates, observes, and guards AI agents end-to-end, correlating quality signals with GPU-level infrastructure data and per-agent token cost ("Tokenomics").

Visit site
09Enterprise

Honeycomb

Observability platform built for AI-era software — agent timeline views and LLM observability sit on top of a fast, OpenTelemetry-native query engine with BubbleUp root-cause analysis.

Visit site

Open Source, Tracing & Evaluation

3 tools

Developer-first, open-source tools focused on instrumentation, trace visualization, and rigorous, code-native evaluation.

10Open Source

Traceloop / OpenLLMetry

Open-source OpenTelemetry instrumentation for LLM applications — add two lines of code and stream traces to Dynatrace, Datadog, Honeycomb, Splunk, New Relic, or any OTel-compatible backend.

Visit site
11Open Source

Phoenix by Arize

Open-source AI development platform for tracing, annotating, and evaluating agents, with an in-app AI engineering agent ("PXI") and one-click self-hosting via terminal, Docker, or Kubernetes.

Visit site
12Open Source

DeepEval

Open-source, pytest-native LLM evaluation framework with 50+ research-backed metrics for hallucination, faithfulness, and agent task completion — runnable in CI/CD or straight from a coding agent's terminal.

Visit site

AI Gateways with Built-in Observability

2 tools

Gateway-layer tools that route and cache LLM calls across providers while logging every request for cost, latency, and quality analysis.

13AI Gateway

Portkey

AI gateway and observability layer for routing, caching, and monitoring LLM calls across 250+ providers, with unified request logs, guardrails, and cost/latency dashboards from a single API.

Visit site
14AI Gateway

Helicone

Open-source AI gateway and LLM observability platform for routing, debugging, and analyzing production AI applications, with request-level tracing, sessions, prompts, and dataset tooling.

Visit site